Please share resumes at Pra...@sapotsystems.com
Job Title: Data Engineer (AWS Glue & Python)
Location: Jersey City, NJ ( Need Only locals)
12+ years
Only H1b
Key Responsibilities
• Over 12+ Years of Experience
• Design, develop, and maintain ETL/ELT pipelines using AWS Glue (PySpark/Scala).
• Build scalable and efficient data integration workflows using AWS services such as S3, Lambda, Glue Crawlers, Glue Studio, Athena, Step Functions, Lake Formation, and Redshift.
• Write high-quality Python and PySpark scripts for complex data transformations.
• Work with structured and unstructured data formats like JSON, Parquet, CSV, Avro, and more.
• Implement data quality checks, metadata management, and logging/monitoring for pipelines.
• Collaborate with analytics, BI, and product teams to understand data requirements.
• Optimize Glue jobs for performance and cost efficiency.
• Manage data security and compliance using IAM, KMS, encryption, and data governance best practices.
• Troubleshoot ETL issues and handle operational support for data workflows.
• Participate in architecture reviews and contribute to data engineering best practices.
Required Skills
• Strong hands-on experience with AWS Glue (Jobs, Crawlers, Workflows, Triggers).
• Proficient in Python and PySpark for large-scale data processing.
• Solid experience working with AWS data ecosystem:
• S3, Athena, Lambda, Step Functions, Redshift, CloudWatch, Lake Formation
• Strong understanding of ETL/ELT concepts, performance tuning, and distributed processing.
• Experience with RDBMS systems like PostgreSQL, MySQL, SQL Server, etc.
• Hands-on experience with data formats: JSON, Parquet, CSV, Avro.
• Good knowledge of version control (Git), CI/CD, and deployment automation.
• Ability to work in an agile and collaborative environment.