About The Role
The DevOps Engineer builds and operates the infrastructure that keeps production services reliable, secure, and available at scale. The role spans cloud infrastructure, Kubernetes, CI/CD, observability, and incident response across development and production environments.
You will partner with software engineers and SREs to reduce operational toil, improve deployment velocity, and establish resilient systems with clear service-level objectives. The team values automation, measurable reliability, and infrastructure that is reproducible through code.
Key Responsibilities
- Design and manage highly available AWS infrastructure using Terraform, including compute, networking, IAM, databases, and monitoring services
- Operate and improve Kubernetes environments, including cluster upgrades, workload scheduling, ingress, autoscaling, secrets management, and service networking
- Build and maintain CI/CD pipelines with GitHub Actions, GitLab CI, or Jenkins for automated testing, security scanning, deployment, and rollback
- Implement observability using Prometheus, Grafana, OpenTelemetry, and centralized logging to track system health, latency, capacity, and service-level objectives
- Automate operational workflows with Python, Go, or Bash, reducing manual intervention across provisioning, deployments, remediation, and environment management
- Participate in on-call rotations and incident response; lead root-cause analysis and deliver durable fixes for production failures
- Partner with engineering teams to define reliability standards, improve release processes, and document runbooks, architecture decisions, and recovery procedures
What We Are Looking For
- 3–8 years of experience in DevOps, SRE, platform engineering, or a closely related infrastructure role
- Hands-on experience operating production workloads in AWS, with strong knowledge of IAM, VPC, EC2, EKS, RDS, S3, and CloudWatch
- Proficiency with Kubernetes and infrastructure as code, particularly Terraform; experience designing reusable modules and managing environments through version control
- Strong understanding of Linux systems, networking, DNS, TLS, containers, and common production troubleshooting practices
- Experience building CI/CD pipelines and implementing deployment strategies such as blue-green, canary, and rolling releases
- Working knowledge of observability, incident management, SLOs, SLIs, and error budgets; strong scripting skills in Python, Go, or Bash
- Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent practical experience; Bonus: experience with Helm, Argo CD, service meshes, policy-as-code, FinOps, or multi-cloud environments