Vacancy catalog
EPAM
Open roleNew

Senior Site Reliability Engineer

EPAMArgentina; Brazil; Chile; Colombia; Mexico
Work model
Hybrid
Experience
3+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
20 tags

Technology context

20

Parsed from the vacancy text; ordered by relevance to this role.

AWSAzurePythonGCPKubernetesCloudDevOpsDockerTerraformLinuxRustSRECI/CDPrometheusAutomationGrafanaBashDNSInsuranceRetail

Full listing

Role description

We are seeking a capable and forward-thinking Senior Site Reliability Engineer (SRE) to join our engineering organization.

In this role, you will serve as the connection point between software development and systems operations, applying engineering principles to automate operational work, expand infrastructure capacity, and keep our systems reliable, resilient, and performing well. Your objective is to build, operate, and safeguard the production environments that support our applications, reducing downtime while enabling fast and secure software releases.

Responsibilities

  • Architect, construct, and maintain cloud infrastructure through modern Infrastructure as Code (IaC) approaches such as Terraform or CloudFormation
  • Develop and refine CI/CD pipelines to streamline software releases, configuration management, and routine operational activities
  • Establish comprehensive logging, monitoring, and alerting frameworks using platforms such as Prometheus, Grafana, or Datadog
  • Define clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to gauge system reliability
  • Address production incidents by leading troubleshooting efforts aimed at restoring service as quickly as possible
  • Facilitate blameless post-mortem reviews to pinpoint root causes and prevent future occurrences
  • Work alongside software developers to fine-tune system performance and forecast capacity requirements
  • Guarantee that services scale properly to accommodate growth and sudden spikes in traffic

Requirements

  • At least 3 years of relevant professional experience
  • Solid foundation in systems administration, DevOps, or systems-oriented software engineering
  • Competency in at least one scripting or programming language, such as Python, Bash, Go, or Rust
  • Practical experience with public cloud platforms, including AWS, Azure, or GCP
  • Direct experience using containerization technologies such as Docker and Kubernetes
  • Thorough knowledge of Linux/Unix system administration and essential networking concepts, including TCP/IP, DNS, HTTP, and SSL/TLS
  • Strong enthusiasm for automation, minimizing repetitive manual work, and designing systems that degrade gracefully under failure
  • Background working within Financial Services, Insurance, or Retail sectors
  • Strong written and spoken proficiency in English at a C1 level or higher