- Work model
- Hybrid
- Experience
- 2-5 years
- Employment
- Not specified
- Compensation
- Not disclosed
- Technology signal
- 13 tags
Technology context
13Parsed from the vacancy text; ordered by relevance to this role.
AWSJavaPythonCloud NativeCloudGoDevOpsRustSREPrometheusAutomationInsuranceRetail
Full listing
Role description
We are looking for a visionary SRE Architect to define and scale reliability architecture for cloud-native, distributed platforms. You will set SRE standards across design, observability, and automation on AWS - apply now to help shape how teams build and run resilient services.
Responsibilities
- Design self-healing, highly available, fault-tolerant architectures on AWS
- Define standards for disaster recovery, multi-region failover, and traffic protection patterns
- Establish organization-wide SRE practices for SLOs, SLIs, and error budget governance
- Architect a global observability strategy for logs, metrics, traces, and APM
- Build internal tools and automation to reduce operational toil and speed up delivery
- Lead capacity planning to ensure predictable performance and scalability
- Introduce chaos engineering practices to continuously validate system resilience
- Partner with engineering teams to embed reliability requirements into design patterns
- Create reference architectures and guidance that improve consistency across services
- Review system designs and recommend improvements for resiliency and operability
Requirements
- 3+ years of solution architecture experience for cloud-native or distributed systems
- 7+ years of site reliability engineering experience applying SRE principles in production
- Technical leadership experience guiding architecture decisions and standards across teams
- Architecture design expertise for high availability, fault tolerance, and multi-region resilience on AWS
- Observability platform experience with logging, metrics, tracing, and APM tooling
- Strong software engineering skills in at least one language such as Go, Python, Java, or Rust
- Hands-on AWS experience across services such as EKS or ECS, RDS, serverless, IAM, and networking
- Reliability engineering knowledge of SLOs, SLIs, and error budgets
- Chaos engineering experience with fault injection and resilience validation practices
- Capacity planning skills for forecasting and scaling highly distributed workloads
- Clear communication skills to align stakeholders and drive consistent SRE practices
- Upper-Intermediate English proficiency (B2)
Nice to have
- AWS Solutions Architect certification
- AWS DevOps Engineer certification
- OpenTelemetry implementation experience
- Experience with tools such as Datadog, Dynatrace, New Relic, Prometheus, or ELK Stack