- Work model
- Remote
- Experience
- 8+ years
- Employment
- Full Time
- Compensation
- Not disclosed
- Technology signal
- 13 tags
Technology context
13Parsed from the vacancy text; ordered by relevance to this role.
Full listing
Role description
Svitla Systems Inc. is looking for a Senior Site Reliability Engineer for a full-time position (40 hours per week) in Argentina.
- Bachelor's degree and 8+ years of professional experience handling large-scale production systems.
- Hands-on experience designing and deploying EKS / AKS clusters.
- Understanding of Kubernetes security best practices, including RBAC, network policies, and PodSecurityPolicies.
- Ability to identify and resolve issues related to Kubernetes, networking, storage, and application deployments.
- Experience migrating workloads to Kubernetes.
- Experience with AWS or a comparable cloud provider, with certification.
- Hands-on experience with Ruby, Terraform, and configuration management tools such as Chef, Ansible, or equivalent.
- Excellent knowledge of large-scale web applications and distributed systems.
- Experience with observability tools such as New Relic and Datadog.
- Expertise in problem-solving and analyzing global-scale distributed systems.
- Excellent written and verbal communication skills.
- Critical thinking and a habit of continuously challenging how and why we do things in order to improve.
- Lead the migration of legacy AWS workloads to Amazon EKS (Elastic Kubernetes Service), using Terraform for reproducible infrastructure and Helm for application packaging.
- Provide weekend support for migration activities.
- Develop Python and Bash automation to streamline containerization, secret management (AWS Secrets Manager), and resource tagging.
- Implement deep-trace monitoring with observability tools to maintain visibility during and after the migration.
- Act as the primary point of contact for resolving Kubernetes incidents, including Pod CrashLoopBackOffs, OOMKills, and VPC CNI networking issues.
- Manage the full incident lifecycle, from real-time troubleshooting (including on weekends) to post-mortem analysis, and build automated guardrails against recurrence.
- Work closely with development, QA, and operations teams to ensure seamless collaboration and efficient workflows.
- Coordinate incident, problem, and change management.
- Participate in an on-call rotation for after-hours emergencies.