Design, implement, and maintain CI/CD pipelines for applications and ML services across environments.
Automate infrastructure provisioning and configuration to improve reliability, repeatability, and deployment speed.
Establish monitoring, logging, and alerting practices to improve system observability and incident response.
Ensure secure access controls, secrets management, and environment hygiene across development and production. MLOps & ML Delivery
Build and maintain ML pipelines for training, validation, packaging, and deployment of models using Python-based workflows.
Enable model versioning, reproducibility, and controlled rollouts (e.g., canary/blue-green) for ML services.
Partner with data science teams to productionize models and define operational SLAs for ML endpoints and batch jobs.
Implement automated quality checks for data/model artifacts to reduce regressions and improve release confidence. LLM Enablement
Support deployment patterns for LLM-based services, including scalable inference, prompt/version management, and runtime monitoring.
Collaborate on integrating LLM capabilities into existing platforms with a focus on reliability, latency, and cost awareness. Minimum Qualifications:
Bachelor’s degree or equivalent in Engineering/Technology/Computer Science (BTech/BE or equivalent); Master’s (MTech/MCA/MSc) is acceptable as listed.
3–5 years of experience in DevOps and MLOps-focused delivery for production systems.
Hands-on experience with Python-based ML workflows and operationalizing ML models into services or batch pipelines.
Strong understanding of CI/CD concepts, release management, and environment promotion strategies.
Experience implementing monitoring and operational practices for reliability and troubleshooting in production. Preferred Qualifications:
Experience building and operating end-to-end MLOps pipelines including model packaging, deployment automation, and lifecycle governance.
Practical exposure to LLM solution delivery, including inference deployment, prompt iteration workflows, and evaluation/monitoring approaches.
Familiarity with containerization and orchestration for ML workloads, and optimizing deployments for performance and scalability.
Experience with infrastructure automation and configuration management to support repeatable ML environments.
Proven ability to collaborate across data science and engineering teams, translating experimentation needs into production-grade systems. Good to have skills: Kubernetes, Docker, Terraform, MLflow, Apache Airflow
A free Jobstore account is required to proceed to the employer site.