About The Role
The role owns the reliability, scalability, and security of core cloud infrastructure supporting high-volume microservices and real-time data pipelines.
The team works closely with backend and platform engineers to build robust CI/CD pipelines, automated monitoring systems, and resilient distributed environments.
Key Responsibilities
- Design, provision, and manage production cloud infrastructure on AWS or GCP using Infrastructure as Code tools such as Terraform
- Build and maintain automated CI/CD deployment pipelines using GitHub Actions, GitLab CI, or ArgoCD to streamline releases
- Manage Kubernetes clusters in production, ensuring optimal resource utilization, auto-scaling, and high availability
- Implement comprehensive observability stacks using Prometheus, Grafana, and Datadog for real-time monitoring and incident alerting
- Conduct security audits, manage IAM policies, and enforce compliance standards across all infrastructure layers
- Participate in an on-call rotation to troubleshoot and resolve infrastructure incidents, performing root cause analysis for system failures
What We Are Looking For
- 3–6 years of experience in DevOps, Site Reliability Engineering, or a closely related infrastructure role
- Strong proficiency in Infrastructure as Code (Terraform, CloudFormation) and configuration management tools
- Hands-on experience with containerization technologies and orchestration platforms, specifically Docker and Kubernetes
- Solid scripting skills in Python, Bash, or Go for automation and tooling development
- Deep understanding of networking fundamentals, TCP/IP, DNS, TLS, and VPC architecture in cloud environments
- Bonus: Experience with service mesh technologies (Istio, Linkerd), eBPF observability tools, or FinOps cost optimization practices