- Work model
- Remote
- Experience
- 3+ years
- Employment
- Not specified
- Compensation
- Not disclosed
- Technology signal
- 25 tags
Technology context
25Parsed from the vacancy text; ordered by relevance to this role.
Full listing
Role description
We are looking for a Senior Data DevOps Engineer to join our team. We are building an Enterprise AI Gateway entirely from the ground up. It acts as the sole entry point that every team throughout the company uses to interact with large language models. It also functions as our means of governing AI at scale, addressing who can access which models, what expenses are incurred, and what gets tracked. The intent is for the system to run itself. Bringing teams, agents, and MCP servers on board, along with managing keys, permissions, limits, and guardrails, should occur through automation instead of manual ticketing. Most of this has yet to be built, giving you the chance to shape it from the earliest design decisions onward. Because the team is small and focused, decisions are made swiftly and your contributions will be easy to see. We are looking for someone who turns to automation by default and who cares about AI governance just as much as about the models themselves.
Responsibilities
- Create and set up the fundamental structure of the Enterprise AI Gateway starting from the ground up
- Develop automation to onboard teams, agents, and MCP servers, including managing keys, permissions, limits, and guardrails
- Build and maintain infrastructure as code to deliver scalable, self-service functionality across the platform
- Set up monitoring, logging, and cost-tracking capabilities to keep AI usage visible and under control
- Design guardrails and governance structures that determine which teams and models can access particular resources
- Maintain the reliability, uptime, and performance of production systems that support the gateway
- Interact directly with platform users, addressing questions and assisting with troubleshooting as issues emerge
- Steadily improve automation to reduce manual work and reliance on ticket-based processes
- Evaluate and adopt new GenAI and agentic AI patterns, frameworks, and protocols within the platform
- Contribute to significant architectural and design decisions as the platform develops past its initial stages
Requirements
- At least 3 years of relevant experience
- Strong Python capabilities used for automation, extensions, and integrations
- Solid SRE background, with proven experience keeping production systems stable
- Proficient with Git for version control
- Experience working with Google Cloud Platform
- Familiarity with LLMOps practices
- Experience using Terraform and Helm for infrastructure automation
- Working knowledge of GenAI/Agentic AI concepts, including relevant patterns, frameworks, and protocols
- Strong communication skills, with the ability to clearly explain technical topics, given regular interaction with platform users to address questions and resolve issues
- Excellent English proficiency (B2 level or higher)
Nice to have
- Practical experience with Google Vertex AI, especially endpoints for model serving and Model Armor
- GCP experience involving services such as BigQuery, Cloud Run, and IAM
- Experience building AI agents, for example using Google's Agent Development Kit (ADK)
- Experience with AWS Bedrock
- Extensive Kubernetes experience, ideally with GKE
- Familiarity with Groovy
- Experience with CI/CD pipelines using Jenkins
- Production experience with an AI gateway, such as LiteLLM or EPAM DIAL, viewed as highly valuable