Long-running vacancy
This listing is older than 30 days but its source has not removed it. Verify availability on the original company page before applying.
- Work model
- Remote
- Experience
- 5+ years
- Employment
- Full Time
- Compensation
- Not disclosed
- Technology signal
- 20 tags
Technology context
20Parsed from the vacancy text; ordered by relevance to this role.
Full listing
Role description
We're looking for a Senior AI/DevOps Engineer who thrives at the intersection of infrastructure, automation, and applied AI. You'll be the person who makes AI-powered features actually ship, scale, and stay healthy in production - building the pipelines, platforms, and guardrails that let the rest of the team move fast with LLMs and ML models. This is a hybrid role: part platform/DevOps engineer, part MLOps practitioner, part AI integrator.
This is a full-time, remote role, with flexibility to work from home. This role works across a globally distributed team spanning India and both US coasts, so comfort with asynchronous collaboration and overlapping across time zones is important. Expect some evening hours to overlap with global teams.
- Design and operate the infrastructure that powers our AI features: LLM integrations, RAG pipelines, vector stores, and model-serving endpoints.
- Build and maintain CI/CD pipelines for both traditional services and AI/ML workloads (model deployment, evaluation, rollback).
- Own observability and monitoring for AI systems - latency, cost, token usage, output quality drift, and failure modes unique to LLM-based features.
- Architect and manage cloud infrastructure on Azure (compute, networking, storage, IAM) using infrastructure-as-code.
- Containerize and orchestrate services with Docker and Kubernetes.
- Integrate with AI APIs and platforms (OpenAI, Anthropic, Azure AI/OpenAI Service, or similar) and build internal tooling to make that integration repeatable across teams.
- Implement guardrails, rate limiting, caching, and cost controls for LLM-powered features.
- Collaborate with product, data, and application engineering teams to translate AI-feature designs into reliable, production-grade systems.
- Coordinate effectively with teammates across India, East Coast, and West Coast time zones.
- Mentor other engineers and help set technical direction for AI infrastructure practices.
- Bachelor's Degree in CS or Engineering.
- 5+ years of professional experience in DevOps, platform engineering, or SRE roles.
- Hands-on experience deploying and operating systems in production.
- Expert level in Python, and proficient in Java (specifically targeting backend product engineering context).
- Strong experience with Azure cloud infrastructure and infrastructure-as-code (Terraform, Bicep, ARM, or similar).
- Solid grasp of CI/CD pipeline design and automation (Azure DevOps, GitHub Actions, or similar).
- Proficiency with Docker and Kubernetes for containerization and orchestration.
- Experience integrating with AI/LLM APIs (OpenAI, Anthropic, Azure AI, or similar) and understanding of RAG architectures, vector databases, and prompt/response evaluation.
- Comfort with monitoring and observability tooling (Prometheus, Grafana, Application Insights, or similar), including metrics specific to AI systems (cost, latency, quality).
- Comfortable working independently and owning problems end-to-end.
- Strong written communication and async collaboration skills, given the distributed, multi-time-zone team.
- Expertise in some of the following frameworks: FastAPI, Pydantic.AI, Pydantic 2, Polars, Pandas, Spring Boot, Hibernate, JDBC.
- Experience with LLM evaluation frameworks and prompt engineering.
- Security/compliance experience around AI systems (data privacy, PII handling, model access controls).
- Background in SOLID principles, clean architecture, and TDD.
- Experience in retail, pricing, or data-intensive SaaS products is a plus.