Vacancy catalog
EPAM
Open role

Senior ML Infrastructure Engineer (DevOps)

EPAMArmenia; Georgia; Kazakhstan; Kyrgyzstan; Uzbekistan
Work model
Remote
Experience
5+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
16 tags

Technology context

16

Parsed from the vacancy text; ordered by relevance to this role.

Node.jsAIAWSMachine LearningPythonKubernetesDevOpsDockerCUDALinuxDatabricksCI/CDAirflowGitAnthropic Claude CodeData DevOps

Full listing

Role description

We are looking for a Senior ML Infrastructure Engineer who enjoys running systems in production. You will work side by side with our applied scientists: they build the models, and you own everything needed to run them reliably for millions of customers. Our team owns its infrastructure end to end, deployment and operations included, so your work is visible, and the ownership is real. Experience with 24x7 on-call rotations for high-load online services in production is required for this role. You will thrive here if you have been the person who gets paged and fixes things, rather than handing deployment and operations to a separate team.

Responsibilities

  • Keep real-time inference services healthy through alerting, live debugging and fixing of production incidents, rollbacks and postmortems, as part of a shared 24x7 on-call rotation
  • Build and improve AWS infrastructure, including infrastructure as code, CI/CD pipelines, Kubernetes, Docker, autoscaling, monitoring, load testing and cost control
  • Care for the GPU fleet through NVIDIA driver and CUDA upgrades, node provisioning and debugging, and capacity management
  • Automate user access, including SSH credentials, service accounts and developer environments for the science team
  • Run real-time production data pipelines, including streaming ingestion into an online feature store and serving features to low-latency endpoints
  • Orchestrate batch jobs with Airflow or Databricks Workflows
  • Partner with applied scientists to productionize their prototypes and keep them healthy
  • Write and review production Python, including validating AI-generated code

Requirements

  • 24x7 on-call experience with high-load online services in production, including real incidents personally debugged and fixed live
  • Strong Linux systems skills, with hands-on experience in Kubernetes, Docker and infrastructure as code
  • Solid experience in Python, Git and CI/CD
  • Experience with workflow orchestration tools such as Airflow, Databricks Workflows or equivalent
  • Proficiency in Apache Airflow, CI/CD and DevOps practices
  • Familiarity with Gen AI in SDLC
  • English proficiency at B2 level or higher

Nice to have

  • Familiarity with Anthropic Claude Code