- Work model
- Remote
- Experience
- 5+ years
- Employment
- Not specified
- Compensation
- Not disclosed
- Technology signal
- 16 tags
Technology context
16Parsed from the vacancy text; ordered by relevance to this role.
AILLMAWSMachine LearningAzurePythonRAGGCPKubernetesCloud NativeCloudDevOpsDockerSRECI/CDAutomation
Full listing
Role description
We're looking for a Senior Site Reliability Engineer (SRE) to join our global delivery team in a fully remote working mode. In this role, you will support a leading global financial market infrastructure provider by ensuring the reliability, scalability, and efficiency of mission-critical platforms used across international markets. You will work closely with development and operations teams to implement SRE best practices, automate operations, and deliver secure, highly available systems. This is an opportunity to have a significant impact on large-scale distributed systems and drive continuous improvement across global technology initiatives.
Responsibilities
- Collaborate with development, security, quality, and operations teams to implement SRE principles
- Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets
- Apply practices to minimize toil, improve resiliency, and optimize capacity
- Troubleshoot and resolve infrastructure and application issues to maintain system availability
- Design, implement, and maintain robust monitoring and alerting systems for applications and infrastructure
- Contribute to automation strategies and CI/CD pipelines to accelerate deployments
- Ensure adherence to security and compliance requirements across cloud-native environments
Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent experience
- Proven experience working on SRE or DevOps in production environments
- Expertise with any cloud platform (AWS, GCP, or Azure)
- Solid understanding of SRE practices, including monitoring, incident response, error budgets, and postmortems
- Proficiency in Python or another scripting/programming language
- Hands-on knowledge of CI/CD pipelines, infrastructure as code, and configuration management tools
- Experience with containerization and orchestration tools (Kubernetes, Docker)
- Strong background in monitoring frameworks for infrastructure and application performance
Nice to have
- Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions
- Cloud and Kubernetes certifications (AWS, GCP, Azure Certified, CKA/CKAD)
- Familiarity with AI/ML model lifecycle management in production environments
- Exposure to advanced automation and optimization of AI-driven applications