I
AWS Glue, Python
I
- Posted 11 hours ago
- Be among the first 10 applicants
Job Description
- Primary skills :AWS Glue-Technology->Cloud Platform->Amazon Webservices DevOps,Technology->Cloud Platform->AWS Data Analytics->AWS Glue DataBrew,Technology->OpenSystem->Python - OpenSystem
- Design, build, and maintain ETL/ELT pipelines using AWS Glue and Python for batch and incremental data processing.
- Develop and optimize Glue Jobs (PySpark/Python) including job parameters, bookmarks, retries, and performance tuning.
- Implement data ingestion, transformation, and validation logic to ensure accuracy, completeness, and consistency of datasets.
- Integrate pipelines with AWS services (e.g., S3, IAM, CloudWatch) to enable secure, observable, and scalable workflows.
- Troubleshoot job failures, analyze logs/metrics, and implement fixes to improve stability and runtime efficiency.
- Collaborate with cross-functional teams to gather requirements, define data mappings, and deliver well-documented solutions.
- Follow engineering best practices including code reviews, version control, and reusable modular coding patterns. Minimum Qualifications:
- Bachelor's degree or equivalent (e.g., BE/BTech/MSc/MCA/MTech).
- 3–5 years of experience in data engineering, ETL development, or data integration roles.
- Strong hands-on experience with AWS Glue and Python for building production-grade data pipelines.
- Working knowledge of core AWS concepts including security basics (IAM), storage patterns, and monitoring.
- Ability to debug data pipeline issues and deliver reliable solutions with clear documentation. Preferred Qualifications:
- Experience with PySpark and distributed data processing patterns within AWS Glue.
- Strong SQL skills and experience working with structured/semi-structured datasets (CSV/JSON/Parquet).
- Exposure to orchestration and scheduling patterns for ETL workflows and dependency management.
- Familiarity with data quality checks, schema evolution handling, and building resilient pipelines.
- Experience collaborating in Agile teams and contributing to CI/CD or automated deployment practices for data jobs. Good to have skills: PySpark, Amazon S3, AWS IAM, Amazon CloudWatch, SQL




