Vacancy catalog
EPAM
Open role

Senior Site Reliability Engineer (SRE)

EPAMArmenia; Georgia; Kazakhstan; Kyrgyzstan; Uzbekistan
Work model
Remote
Experience
3+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
18 tags

Technology context

18

Parsed from the vacancy text; ordered by relevance to this role.

AWSPythonKubernetesCloudReliability EngineeringDevOpsAPITerraformSRECI/CDPrometheusGrafanaPerformance OptimizationBashAIOpsLoad TestingLog management toolsStress testing

Full listing

Role description

We are looking for a Senior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis.

Responsibilities

  • Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooks
  • Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality
  • Design and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanisms
  • Apply a performance engineering mindset through load testing, capacity analysis, and latency profiling
  • Automate operational toil through scripting and infrastructure-as-code
  • Accelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effort
  • Lead incident response practices including on-call readiness and blameless post-mortems
  • Collaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements

Requirements

  • 3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems
  • Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events
  • Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews
  • Experience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layers
  • Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelines
  • Background in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotations
  • A proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teams
  • Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations
  • English Level: B2+ (Upper-Intermediate) or higher

Nice to have

  • Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively
  • Skills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks
  • Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation
  • Experience with Datadog or similar enterprise observability platforms
  • Background in evangelizing best practices and setting standards across engineering teams
  • Exposure to programmatic advertising or adtech platforms