Vacancy catalog
EPAM
Open roleUpdated

Senior Site Reliability Engineer (SRE)

EPAMSpain
Work model
Remote
Experience
5+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
18 tags

Technology context

18

Parsed from the vacancy text; ordered by relevance to this role.

AILLMAWSMachine LearningAzurePythonRAGGCPKubernetesCloudDevOpsDockerTerraformJenkinsSREAnsibleCI/CDAutomation

Full listing

Role description

We're looking for a Senior Site Reliability Engineer (SRE) to join our team in Spain in a remote working mode. In this role, you will collaborate with development, operations, security and quality teams to ensure highly reliable, scalable and efficient systems for business-critical applications in the financial domain. You will focus on implementing SRE practices, reducing toil through automation and driving operational excellence while meeting strict Service Level Objectives (SLOs).

This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences.

Responsibilities

  • Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services
  • Collaborate with cross-functional teams to embed reliability into application and infrastructure design
  • Automate operational tasks to reduce manual toil and improve service performance
  • Troubleshoot and resolve infrastructure and application incidents quickly and effectively
  • Implement robust monitoring and observability systems to detect and prevent outages
  • Plan capacity and scaling strategies to ensure high availability and resiliency
  • Contribute to incident postmortems and continuous improvement initiatives
  • Support the adoption of SRE best practices across all SDLC stages

Requirements

  • Bachelor's degree in Computer Science, Engineering or related field
  • Proven experience working in cloud environments (AWS, GCP or Azure)
  • Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation)
  • Proficiency in Python or other scripting language for automation tasks
  • Strong understanding of monitoring tools and observability frameworks
  • Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab)
  • Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes

Nice to have

  • Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions
  • Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies
  • Background in DevOps practices and agile delivery frameworks
  • Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments