Vacancy catalog
EPAM
Open role

Senior AI Reliability Engineer

EPAMUkraine
Work model
Remote
Experience
4+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
22 tags

Technology context

22

Parsed from the vacancy text; ordered by relevance to this role.

AI AgentsAILLMAWSGenAIMachine LearningAzurePythonRAGGCPKubernetesCloudDevOpsAPITerraformDatabricksSRECI/CDSplunkAutomationGrafanaAI Solution Engineering

Full listing

Role description

EPAM's Operational Intelligence practice is building a new capability: AI Reliability Engineering (AIRE) - applying SRE principles and cloud-native practices to the lifecycle of production AI/ML and LLM systems. As clients move GenAI applications and agentic systems from pilot into production, they discover that classic APM tells them nothing about token latency, cost per request, semantic drift or hallucinations. That gap is what this role closes.

You will instrument, monitor and harden production AI systems, define AI-native service level objectives and build the accelerators and reference implementations our practice reuses across accounts. This is an engineering role, not an L1/L2 support role - no 24/7 on-call rotation.

What You'll Get

  • A genuinely new discipline inside EPAM - you define how it is done, not follow an existing runbook
  • Engineering focus without 24/7 on-call rotation
  • Funded certification and enablement tracks (Anthropic/Claude, Databricks, AI & Data Observability learning paths)
  • Cross-client exposure and a direct path into presales and solution engineering

Responsibilities

  • Instrument production LLM, RAG and agentic applications with AI telemetry based on OpenTelemetry and APM-native AI monitoring
  • Implement distributed tracing across multi-model chains, agent workflows and retrieval-augmented generation pipelines to profile systemic latency and failure points
  • Define and measure AI-native SLIs and SLOs: time to first token (TTFT), throughput, error and refusal rates, cost per request, semantic drift, hallucination boundaries, contextual accuracy
  • Set up structured semantic logging and prompt/response monitoring for quality analysis
  • Build and run evaluation loops for output quality and safety - golden sets, LLM-as-a-judge, Ragas/DeepEval-style frameworks - and wire them into CI/CD and runtime
  • Track token-based cloud spend, model API rate limits and quota consumption; drive AI cost optimization
  • Configure AI gateways for API load balancing, failover and fallback models across multiple LLM providers
  • Implement guardrails: prompt injection and jailbreak filtering, output compliance, bias and safety constraints
  • Design detection, triage, restore and problem management workflows for AI incidents; integrate autonomous AI agents into RCA to parse logs, form hypotheses and correlate state changes
  • Support rollback, canary and fail-safe patterns for model, prompt and configuration releases; maintain reproducibility through versioning of data, code, prompts and models
  • Build practice accelerators, reference architectures and internal enablement materials; support presales and client assessments

Requirements

  • 4+ years in SRE, DevOps, platform or observability engineering, including hands-on work with production AI/ML or LLM workloads
  • Solid SRE fundamentals: Golden Signals, SLI/SLO definition, error budgets and burn rate, incident lifecycle, ITIL basics
  • Strong Python for instrumentation, automation and evaluation tooling
  • Practical experience with OpenTelemetry and at least one APM/observability platform: New Relic, Datadog, Grafana LGTM stack, Splunk or Elastic
  • Hands-on production experience with at least one cloud platform (Azure preferred, AWS or GCP) and Kubernetes
  • Working understanding of LLM application architecture: prompts, embeddings and vector stores, RAG, agent orchestration (LangChain / LangGraph or equivalent)
  • MLOps awareness: model lifecycle (training vs. inference), model endpoints, containerization, deployment and rollback patterns
  • Infrastructure as Code with Terraform; CI/CD experience (Azure DevOps, GitLab CI, GitHub Actions)
  • B2+ English - the role is client-facing and requires clear written and spoken technical communication

Nice to have

  • Experience with AI-specific observability and evaluation tooling: Traceloop/OpenLLMetry, Langfuse, Arize Phoenix, Ragas, DeepEval, MLflow
  • Distributed inference serving at scale: vLLM, KServe, Ray Serve, Kubernetes-native LLM orchestration, GPU capacity planning
  • AI security: OWASP LLM Top 10, prompt injection defense, guardrail frameworks (NeMo Guardrails, Llama Guard)
  • Databricks (incl. Mosaic AI / MLflow) or Azure AI Foundry experience
  • Certifications: Anthropic Claude, Azure AI Engineer, AWS ML Specialty, Databricks GenAI
  • FinOps for AI workloads - token and GPU cost modeling
  • Data Reliability Engineering background (data quality, pipeline SLOs) - AI reliability starts with data reliability
  • Mentoring or team lead experience