Vacancy catalog
EPAM
Open roleNew

Site Reliability Engineer

EPAMMexico, Morelia
Work model
Hybrid
Experience
2+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
20 tags

Technology context

20

Parsed from the vacancy text; ordered by relevance to this role.

AWSAzurePythonGCPKubernetesCloudDevOpsDockerTerraformLinuxRustSRECI/CDPrometheusAutomationGrafanaBashDNSInsuranceRetail

Full listing

Role description

We are seeking a proactive Site Reliability Engineer to strengthen the reliability, scalability, and safety of production environments. You will bridge software development and operations through automation, observability, and incident response - apply now to help reduce downtime and enable fast, safe releases.

Responsibilities

  • Design and maintain cloud infrastructure using Infrastructure as Code practices
  • Build and optimize CI/CD pipelines to automate deployments and operational workflows
  • Implement logging, monitoring, and alerting to improve observability and reliability
  • Define and track Service Level Objectives and Service Level Indicators with clear reporting
  • Respond to production incidents and drive rapid service restoration
  • Lead blameless post-mortems to identify root causes and prevent recurrence
  • Partner with engineers to improve performance, scalability, and capacity planning
  • Automate repetitive operational tasks to reduce toil and operational risk
  • Harden production environments to improve resilience and safe change practices

Requirements

  • 2+ years of experience in site reliability engineering, DevOps, or systems administration
  • Hands-on experience with Infrastructure as Code using Terraform or CloudFormation
  • Hands-on experience building and improving CI/CD pipelines for automated deployments
  • Strong troubleshooting and incident response leadership skills in production environments
  • Solid project skills to coordinate reliability work with software development teams
  • Proficiency in scripting or programming with Python, Bash, Go, or Rust
  • Cloud platform experience with AWS, Azure, or GCP
  • Containerization experience with Docker and Kubernetes
  • Deep Linux/Unix administration knowledge and networking fundamentals (TCP/IP, DNS, HTTP, SSL/TLS)
  • Strong communication skills with a reliability mindset focused on automation and reducing toil
  • Advanced English proficiency (C1, Advanced)

Nice to have

  • Experience with Prometheus, Grafana, or Datadog