About The Role
The role focuses on building, scaling, and maintaining the infrastructure that powers a high-throughput, low-latency cloud platform. You will design self-healing systems, optimize continuous deployment pipelines, and ensure maximum uptime and performance across multi-region Kubernetes clusters.
Working closely with software engineering teams, the role acts as a bridge between application code and underlying infrastructure. The focus is on automating everything, eliminating manual toil, and treating infrastructure strictly as code to support rapid, reliable software delivery.
Key Responsibilities
- Design, provision, and manage multi-region cloud infrastructure on AWS using Terraform for declarative Infrastructure as Code (IaC)
- Maintain and optimize production Kubernetes clusters (EKS), focusing on network policies, autoscaling behaviors, and resource efficiency
- Build and maintain robust CI/CD pipelines using GitHub Actions or GitLab CI to automate testing, security scanning, and containerized deployments
- Implement comprehensive monitoring, logging, and distributed tracing systems using Prometheus, Grafana, and the ELK stack to establish clear SLOs/SLIs
- Participate in a shared blameless on-call rotation, leading incident response, conducting root-cause analysis (RCA), and implementing preventative automation
- Collaborate with security teams to enforce IAM least-privilege policies, secret management using HashiCorp Vault, and vulnerability scanning in build pipelines
What We Are Looking For
- 3–6 years of experience in a DevOps, SRE, or Infrastructure Engineering role supporting production environments at scale
- Strong proficiency with Infrastructure as Code, specifically writing modular, reusable Terraform
- Production-level experience orchestrating containerized workloads with Kubernetes, including Helm chart management
- Solid scripting and automation skills in Python, Go, or Bash, with a strong emphasis on writing clean, version-controlled code
- Deep understanding of Linux systems administration, TCP/IP networking, DNS, and load balancing topologies
- Bonus: Experience with service meshes (Istio/Linkerd), GitOps deployment patterns (ArgoCD), or managing high-volume PostgreSQL/Kafka deployments