PYTHON DATA PROCESSING, DATAPIPELINE, ETL Pandas, NumPy, Apache Airflow, PySpark, Data Quality Frameworks
Key Responsibilities:
- Design, develop, and maintain Python-based data processing workflows for structured and semi-structured datasets.
- Build and optimize ETL processes to ingest, transform, validate, and publish data for downstream consumption.
- Develop and manage data pipelines, ensuring reliability, scalability, and performance across workloads.
- Write efficient SQL queries for data extraction, transformation, reconciliation, and reporting needs.
- Implement data quality checks, validation rules, and monitoring to ensure accuracy and completeness of datasets.
- Troubleshoot pipeline failures, identify root causes, and implement preventive fixes to reduce recurrence.
- Collaborate with cross-functional teams to gather requirements, define data contracts, and deliver well-documented solutions.
- Maintain clear technical documentation for pipeline logic, transformations, and operational runbooks.
- Follow best practices for code quality, version control, testing, and peer reviews to ensure maintainable delivery. Minimum Qualifications:
- BTECH, MTECH, MCA, or MSC in Computer Science, Information Technology, Data/Software Engineering, or a related field.
- 3–5 years of experience in Python development focused on data processing and transformation use cases.
- Hands-on experience building ETL workflows and maintaining production-grade data pipelines.
- Strong SQL skills with experience in writing complex queries and optimizing performance.
- Practical understanding of data processing concepts such as batching, partitioning, schema handling, and data validation. Preferred Qualifications:
- Experience designing scalable pipeline patterns (incremental loads, CDC-style approaches, idempotent processing) and improving pipeline reliability.
- Strong debugging skills with the ability to analyze data issues across multiple pipeline stages and resolve them efficiently.
- Familiarity with orchestration and scheduling concepts for ETL jobs and dependency management.
- Exposure to performance tuning for Python and SQL workloads, including memory-efficient processing and query optimization.
- Proven ability to work with stakeholders to translate data requirements into robust, well-tested implementations.