Vacancy catalog
EPAM
Open roleUpdated

Senior Site Reliability Engineer (SRE)

EPAMPortugal
Work model
Remote
Experience
5+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
16 tags

Technology context

16

Parsed from the vacancy text; ordered by relevance to this role.

AILLMAWSMachine LearningAzurePythonRAGGCPKubernetesCloud NativeCloudDevOpsDockerSRECI/CDAutomation

Full listing

Role description

We're looking for a Senior Site Reliability Engineer (SRE) to join our global delivery team in a fully remote working mode. In this role, you will support a leading global financial market infrastructure provider by ensuring the reliability, scalability, and efficiency of mission-critical platforms used across international markets. You will work closely with development and operations teams to implement SRE best practices, automate operations, and deliver secure, highly available systems. This is an opportunity to have a significant impact on large-scale distributed systems and drive continuous improvement across global technology initiatives.

Responsibilities

  • Collaborate with development, security, quality, and operations teams to implement SRE principles
  • Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets
  • Apply practices to minimize toil, improve resiliency, and optimize capacity
  • Troubleshoot and resolve infrastructure and application issues to maintain system availability
  • Design, implement, and maintain robust monitoring and alerting systems for applications and infrastructure
  • Contribute to automation strategies and CI/CD pipelines to accelerate deployments
  • Ensure adherence to security and compliance requirements across cloud-native environments

Requirements

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience
  • Proven experience working on SRE or DevOps in production environments
  • Expertise with any cloud platform (AWS, GCP, or Azure)
  • Solid understanding of SRE practices, including monitoring, incident response, error budgets, and postmortems
  • Proficiency in Python or another scripting/programming language
  • Hands-on knowledge of CI/CD pipelines, infrastructure as code, and configuration management tools
  • Experience with containerization and orchestration tools (Kubernetes, Docker)
  • Strong background in monitoring frameworks for infrastructure and application performance

Nice to have

  • Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions
  • Cloud and Kubernetes certifications (AWS, GCP, Azure Certified, CKA/CKAD)
  • Familiarity with AI/ML model lifecycle management in production environments
  • Exposure to advanced automation and optimization of AI-driven applications