Job Overview
Role Summary:
We are seeking a talented Site Reliability Engineer (SRE) to ensure high reliability, observability, and scalability of multi-tenant cloud applications on Google Cloud Platform (GCP) and AWS.
Responsibilities:
• Manage and scale production Kubernetes (GKE / EKS) clusters running mission-critical workloads.
• Build and maintain infrastructure-as-code using Terraform and GitOps practices with ArgoCD.
• Implement comprehensive observability, logging, and alerting systems using Prometheus, Grafana, and Datadog.
• Lead incident response triage, conduct root cause analyses (RCA), and implement preventive remedies.
• Automate routine operational procedures and build self-healing infrastructure.
Requirements:
• 4+ years of professional DevOps or SRE experience managing production cloud platforms.
• Deep understanding of Kubernetes networking, ingress controllers, helm charts, and container security.
• Strong automation scripting skills in Python, Bash, or Go.
• Experience defining and monitoring Service Level Objectives (SLOs) and Error Budgets.
