Vacancy catalog
EPAM
Open role

Senior Data Reliability Engineer

EPAMUkraine
Work model
Remote
Experience
4+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
22 tags

Technology context

22

Parsed from the vacancy text; ordered by relevance to this role.

AIAWSMachine LearningAzurePythonGCPCloudDevOpsSparkTerraformSnowflakeDatabricksSREAnsibleCI/CDSplunkAutomationGrafanaSQLAirflowPower BIData Software Engineering

Full listing

Role description

EPAM's Operational Intelligence practice is expanding from classic infrastructure and application observability into Data & AI Reliability Engineering. We are looking for a Data Reliability Engineer who will apply SRE principles to data pipelines and data platforms - moving our clients from reactive firefighting to proactive management of data quality, so that data stays accurate, complete, fresh and available end-to-end.

This is an engineering role, not an L1/L2 support role: you will build detection, automation and prevention rather than sit in a 24/7 on-call rotation. You will work across client engagements and practice-level initiatives - reusable data quality frameworks, accelerators, and internal enablement.

What You'll Get

  • A practice that is actively building a new capability - your work becomes the standard, not a copy of someone else's runbook
  • Engineering focus without 24/7 on-call rotation
  • Internal certification and enablement tracks (Databricks, AI/LLM, Data Observability learning path in Learn)
  • Cross-client exposure: enterprise-scale retail, financial services and manufacturing accounts

Responsibilities

  • Design and implement end-to-end data observability across pipelines and data platforms, covering the core data quality pillars: freshness, volume, schema, completeness and accuracy
  • Embed automated data quality checks into CI/CD pipelines and orchestrators (e.g., dbt tests, Great Expectations, native platform checks)
  • Configure anomaly detection - including dynamic and ML-based thresholds - for data drift, volume anomalies and pipeline failures
  • Define, measure and report Data SLIs, SLOs and error budgets together with business and data product stakeholders
  • Perform root cause analysis by tracing data lineage upstream to the exact origin of a failure; drive problem management so incidents do not recur
  • Design architectural guardrails for data pipelines: circuit breakers, dead-letter queues, retry and rollback mechanisms, self-healing patterns
  • Implement monitoring and alerting as code (Terraform / GitOps) instead of manual UI configuration
  • Reduce alert noise through event correlation, tagging standards and actionable alert design
  • Build reusable data quality and observability frameworks, templates and accelerators adopted by multiple data product teams
  • Lead post-incident reviews and translate learnings into systemic platform improvements
  • Contribute to presales activities, client-facing assessments and internal training materials for the practice

Requirements

  • 4+ years in SRE, DevOps, Data Engineering or Data Platform Operations, with at least 1 - 2 years focused on data platforms or data pipelines
  • Solid SRE fundamentals: Golden Signals, SLI/SLO definition and calculation, error budgets and burn rate, incident lifecycle, ITIL basics
  • Strong SQL and practical Python for automation and validation scripting
  • Hands-on production experience with at least one cloud platform: Azure, AWS or GCP
  • Experience with at least one modern data platform and orchestration stack - Databricks (preferred), Snowflake, Spark, Airflow, Azure Data Factory, dbt
  • Working knowledge of observability tooling: New Relic, Datadog, Splunk, Elastic Stack, Grafana / OpenTelemetry (any two or more)
  • Infrastructure as Code with Terraform (Ansible is a plus) and CI/CD experience (Azure DevOps, GitLab CI, GitHub Actions)
  • Experience with incident and change management tooling: ServiceNow, PagerDuty or equivalent
  • Ability to troubleshoot complex distributed data issues under SLA pressure
  • B2+ English - the role is client-facing and requires clear written and spoken technical communication

Nice to have

  • Databricks certification (Data Engineer Associate / Professional) or equivalent cloud data certification
  • Hands-on experience with dedicated data observability platforms (Monte Carlo, Soda, Anomalo, Great Expectations)
  • Data catalog and lineage tooling (Unity Catalog, OpenLineage, Collibra)
  • FinOps: cloud, platform and telemetry cost optimization
  • Exposure to AI/ML pipeline monitoring or AI Reliability Engineering
  • Power BI or another BI layer, from a monitoring and reliability perspective
  • Event correlation / AIOps experience (New Relic Decisions, IBM NOI, ServiceNow ITOM)
  • Mentoring or team lead experience