Vacancy catalog
Svitla Systems
Open role

SENIOR SITE RELIABILITY ENGINEER

Svitla SystemsArgentina
Work model
Remote
Experience
8+ years
Employment
Full Time
Compensation
Not disclosed
Technology signal
13 tags

Technology context

13

Parsed from the vacancy text; ordered by relevance to this role.

Full listing

Role description

Svitla Systems Inc. is looking for a Senior Site Reliability Engineer for a full-time position (40 hours per week) in Argentina.

  • Bachelor's degree and 8+ years of professional experience handling large-scale production systems.
  • Hands-on experience designing and deploying EKS / AKS clusters.
  • Understanding of Kubernetes security best practices, including RBAC, network policies, and PodSecurityPolicies.
  • Ability to identify and resolve issues related to Kubernetes, networking, storage, and application deployments.
  • Experience migrating workloads to Kubernetes.
  • Experience with AWS or a comparable cloud provider, with certification.
  • Hands-on experience with Ruby, Terraform, and configuration management tools such as Chef, Ansible, or equivalent.
  • Excellent knowledge of large-scale web applications and distributed systems.
  • Experience with observability tools such as New Relic and Datadog.
  • Expertise in problem-solving and analyzing global-scale distributed systems.
  • Excellent written and verbal communication skills.
  • Critical thinking and a habit of continuously challenging how and why we do things in order to improve.
  • Lead the migration of legacy AWS workloads to Amazon EKS (Elastic Kubernetes Service), using Terraform for reproducible infrastructure and Helm for application packaging.
  • Provide weekend support for migration activities.
  • Develop Python and Bash automation to streamline containerization, secret management (AWS Secrets Manager), and resource tagging.
  • Implement deep-trace monitoring with observability tools to maintain visibility during and after the migration.
  • Act as the primary point of contact for resolving Kubernetes incidents, including Pod CrashLoopBackOffs, OOMKills, and VPC CNI networking issues.
  • Manage the full incident lifecycle, from real-time troubleshooting (including on weekends) to post-mortem analysis, and build automated guardrails against recurrence.
  • Work closely with development, QA, and operations teams to ensure seamless collaboration and efficient workflows.
  • Coordinate incident, problem, and change management.
  • Participate in an on-call rotation for after-hours emergencies.