Vacancy catalog
EPAM
Open roleNew

Lead Data DevOps Engineer

EPAMBrazil; Colombia; Mexico
Work model
Remote
Experience
5+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
25 tags

Technology context

25

Parsed from the vacancy text; ordered by relevance to this role.

AI AgentsAIAWSGenAIMCPPythonGCPKubernetesCloudDevOpsTerraformJenkinsSRECI/CDHelmAutomationGitData DevOpsData DevOps AwsGoogle Gemini EnterpriseGoogle Kubernetes EngineGoogle Vertex AIGroovyLiteLLMLLMOps

Full listing

Role description

We are seeking a Lead Data DevOps Engineer to become part of our team. We are developing an Enterprise AI Gateway completely from scratch. It functions as the single access point every team across the company relies on to work with large language models. It also serves as our mechanism for controlling AI at scale, covering who can access which models, associated costs, and what activity gets recorded. The vision is for the system to operate independently. Bringing on teams, agents, and MCP servers, along with handling keys, permissions, limits, and guardrails, should be driven by automation rather than manual ticket handling. Most of this remains unbuilt, so you will have influence over it from the very first design decisions. Given the small, focused nature of the team, decisions happen quickly and your impact will be readily apparent. We are seeking someone who automates as a first instinct and who places as much importance on AI governance as on the models themselves.

Responsibilities

  • Design and establish the core structure of the Enterprise AI Gateway from the very beginning
  • Build automation for onboarding teams, agents, and MCP servers, covering the setup of keys, permissions, limits, and guardrails
  • Develop and sustain infrastructure as code to enable scalable, self-service capabilities across the platform
  • Implement monitoring, logging, and cost-tracking tools to maintain clear oversight of AI usage
  • Create guardrails and governance frameworks that regulate which teams and models can reach specific resources
  • Ensure production systems supporting the gateway remain reliable, available, and performant
  • Engage directly with platform users, answering questions and helping troubleshoot problems as they come up
  • Continuously enhance automation to cut down on manual effort and dependence on ticket-driven processes
  • Review and integrate new GenAI and agentic AI patterns, frameworks, and protocols into the platform
  • Help drive important architectural and design choices as the platform grows beyond its early stages

Requirements

  • A minimum of 5 years of relevant experience
  • At least one year of experience leading and managing teams
  • Strong Python skills applied to automation, extensions, and integrations
  • Solid SRE experience, with a genuine track record of maintaining stable production systems
  • Skilled with Git for version control
  • Background working with Google Cloud Platform
  • Familiarity with LLMOps practices
  • Experience applying Terraform and Helm for infrastructure automation
  • Working knowledge of GenAI/Agentic AI concepts, including related patterns, frameworks, and protocols
  • Excellent communication skills, with the ability to explain technical topics clearly, given daily interaction with platform users to answer questions and resolve issues
  • Excellent English proficiency (B2 level or higher)

Nice to have

  • Practical experience with Google Vertex AI, particularly endpoints for model serving and Model Armor
  • GCP experience covering services such as BigQuery, Cloud Run, and IAM
  • Experience developing AI agents, for example through Google's Agent Development Kit (ADK)
  • Experience with AWS Bedrock
  • Broad Kubernetes experience, ideally involving GKE
  • Familiarity with Groovy
  • Experience with CI/CD pipelines built on Jenkins
  • Production experience with an AI gateway, such as LiteLLM or EPAM DIAL, seen as highly valuable