Vacancy catalog
EPAM
Open roleNew

SRE Architect

EPAMMexico, Guadalajara
Work model
Hybrid
Experience
2-5 years
Employment
Not specified
Compensation
Not disclosed
Technology signal
13 tags

Technology context

13

Parsed from the vacancy text; ordered by relevance to this role.

AWSJavaPythonCloud NativeCloudGoDevOpsRustSREPrometheusAutomationInsuranceRetail

Full listing

Role description

We are looking for a visionary SRE Architect to define and scale reliability architecture for cloud-native, distributed platforms. You will set SRE standards across design, observability, and automation on AWS - apply now to help shape how teams build and run resilient services.

Responsibilities

  • Design self-healing, highly available, fault-tolerant architectures on AWS
  • Define standards for disaster recovery, multi-region failover, and traffic protection patterns
  • Establish organization-wide SRE practices for SLOs, SLIs, and error budget governance
  • Architect a global observability strategy for logs, metrics, traces, and APM
  • Build internal tools and automation to reduce operational toil and speed up delivery
  • Lead capacity planning to ensure predictable performance and scalability
  • Introduce chaos engineering practices to continuously validate system resilience
  • Partner with engineering teams to embed reliability requirements into design patterns
  • Create reference architectures and guidance that improve consistency across services
  • Review system designs and recommend improvements for resiliency and operability

Requirements

  • 3+ years of solution architecture experience for cloud-native or distributed systems
  • 7+ years of site reliability engineering experience applying SRE principles in production
  • Technical leadership experience guiding architecture decisions and standards across teams
  • Architecture design expertise for high availability, fault tolerance, and multi-region resilience on AWS
  • Observability platform experience with logging, metrics, tracing, and APM tooling
  • Strong software engineering skills in at least one language such as Go, Python, Java, or Rust
  • Hands-on AWS experience across services such as EKS or ECS, RDS, serverless, IAM, and networking
  • Reliability engineering knowledge of SLOs, SLIs, and error budgets
  • Chaos engineering experience with fault injection and resilience validation practices
  • Capacity planning skills for forecasting and scaling highly distributed workloads
  • Clear communication skills to align stakeholders and drive consistent SRE practices
  • Upper-Intermediate English proficiency (B2)

Nice to have

  • AWS Solutions Architect certification
  • AWS DevOps Engineer certification
  • OpenTelemetry implementation experience
  • Experience with tools such as Datadog, Dynatrace, New Relic, Prometheus, or ELK Stack