- Work model
- Office
- Experience
- 2+ years
- Employment
- Full Time
- Compensation
- Not disclosed
- Technology signal
- 13 tags
Technology context
13Parsed from the vacancy text; ordered by relevance to this role.
Node.jsAIAWSAI-Assisted DevelopmentPythonKubernetesCloudGoDevOpsTerraformLinuxPrometheusGrafana
Full listing
Role description
About the team
KubeOps operates Wix's production Kubernetes platform at scale: a multi-DC EKS fleet with clusters running up to ~50,000 pods and ~2,000 nodes, hosting thousands of services across the company. We own the platform end-to-end - reliability, scale, upgrades, autoscaling and networking.
Job description
- Master the Cluster: Manage the lifecycle of our large-scale Kubernetes environments, without disrupting the thousands of services running on it.
- Build the Platform: Own and evolve the cluster building blocks - Kubernetes Architecture, Karpenter configurations, addons, networking, autoscaling. Where it makes sense, platformize them so other infra teams can consume them cleanly
- Pioneer AI-Driven Ops: Lead the charge in integrating AI into our ecosystem - building LLM-powered tools that accelerate investigation, planning, and coding for the whole team.
- Full-Stack Ownership: You'll take ambiguous tasks and transform them into high-quality outcomes, owning every decision along the way.
Requirements
- +2 years in managing infrastructure, operating production systems at large scale
- 2+ years running production Kubernetes - EKS preferred with a strong understanding of core control-plane and cluster components, with hands-on experience debugging, upgrading, and tuning autoscaling and cluster networking
- 2+ years with a major cloud provider (AWS preferred) - solid grasp of core compute, networking, and IAM.
- 2+ years with Infrastructure-as-Code, writing and maintaining reusable modules (Terraform, Pulumi, Crossplane, etc.)
- Familiarity with observability at scale - Prometheus, Grafana, alert design
- Experience building internal tooling in Go or Python
- Hands-on experience with AI-assisted development tools and agents (Claude Code, Codex, Cursor) for knowledge gathering, debugging, coding, planning, and design - with the ability to integrate these into concrete workflows that demonstrably accelerate your work.
Advantage
- Cluster autoscaling in production
- Cost / capacity optimization at fleet scale - Karpenter consolidation, bin-packing, spot strategy
- Linux internals and node-level debugging - when a Kubernetes problem turns into a Linux problem, you can take it from there
- Track record of leading projects end-to-end: scoping, execution and delivery.