Vacancy catalog
EPAM
Open role

Senior DevOps Engineer

EPAMArmenia; Georgia
Work model
Remote
Experience
3+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
14 tags

Technology context

14

Parsed from the vacancy text; ordered by relevance to this role.

AIMachine LearningKubernetesCloudHigh AvailabilityReliability EngineeringDevOpsScalabilityTerraformCI/CDHelmAutomationIaC EngineeringKustomize

Full listing

Role description

We are seeking a Senior DevOps Engineer to join an AI Workbench Platform team focused on operationalizing domain foundation models. These are large models trained on log, seismic, drilling, and production data, analogous to general-purpose large language models but specialized for oil & gas subsurface and production domains. While these models exist today and are maturing, there is currently no platform to commercialize them, enable internal teams (geo units, business units, data scientists) to use them at scale, or allow external customers to interactively consume or fine-tune them. In this role, you will help design and build the infrastructure foundation that brings these powerful domain models to production.

Responsibilities

  • Design, build, and maintain scalable infrastructure to support the AI Workbench Platform and its underlying domain foundation models
  • Implement and manage Kubernetes clusters with multi-GPU scheduling capabilities to support large-scale model training and inference workloads
  • Develop and maintain Infrastructure as Code using Terraform to provision cloud resources reliably and reproducibly
  • Package, deploy, and manage applications using Helm and Kustomize across multiple environments
  • Collaborate with data scientists, ML engineers, and business units to enable seamless model consumption and fine-tuning workflows at scale
  • Ensure platform reliability, scalability, and security for both internal teams and external customers
  • Optimize resource utilization and cost efficiency across GPU-intensive workloads
  • Establish CI/CD pipelines and automation to accelerate platform delivery and model deployment
  • Monitor system performance and troubleshoot production issues to maintain high availability
  • Contribute to platform architecture decisions and best practices for MLOps at enterprise scale

Requirements

  • 3+ years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering roles
  • Expertise in Kubernetes with proven experience in multi-GPU scheduling for AI/ML workloads
  • Proficiency in Terraform for Infrastructure as Code and cloud resource management
  • Skills in Helm and Kustomize for Kubernetes application packaging and configuration management
  • Background in building and operating production-grade platforms that support large-scale, distributed workloads
  • Understanding of MLOps principles and infrastructure requirements for training and serving large foundation models
  • Capability to collaborate cross-functionally with data scientists, ML engineers, and business stakeholders
  • Excellent command of written and spoken English (B2+ level)

Nice to have

  • Prior experience with LightOps infrastructure
  • Familiarity with on-premises infrastructure environments
  • Knowledge of High-Performance Computing (HPC) systems and workloads