- Work model
- Remote
- Experience
- 4+ years
- Employment
- Not specified
- Compensation
- Not disclosed
- Technology signal
- 22 tags
Technology context
22Parsed from the vacancy text; ordered by relevance to this role.
Full listing
Role description
EPAM's Operational Intelligence practice is expanding from classic infrastructure and application observability into Data & AI Reliability Engineering. We are looking for a Data Reliability Engineer who will apply SRE principles to data pipelines and data platforms - moving our clients from reactive firefighting to proactive management of data quality, so that data stays accurate, complete, fresh and available end-to-end.
This is an engineering role, not an L1/L2 support role: you will build detection, automation and prevention rather than sit in a 24/7 on-call rotation. You will work across client engagements and practice-level initiatives - reusable data quality frameworks, accelerators, and internal enablement.
What You'll Get
- A practice that is actively building a new capability - your work becomes the standard, not a copy of someone else's runbook
- Engineering focus without 24/7 on-call rotation
- Internal certification and enablement tracks (Databricks, AI/LLM, Data Observability learning path in Learn)
- Cross-client exposure: enterprise-scale retail, financial services and manufacturing accounts
Responsibilities
- Design and implement end-to-end data observability across pipelines and data platforms, covering the core data quality pillars: freshness, volume, schema, completeness and accuracy
- Embed automated data quality checks into CI/CD pipelines and orchestrators (e.g., dbt tests, Great Expectations, native platform checks)
- Configure anomaly detection - including dynamic and ML-based thresholds - for data drift, volume anomalies and pipeline failures
- Define, measure and report Data SLIs, SLOs and error budgets together with business and data product stakeholders
- Perform root cause analysis by tracing data lineage upstream to the exact origin of a failure; drive problem management so incidents do not recur
- Design architectural guardrails for data pipelines: circuit breakers, dead-letter queues, retry and rollback mechanisms, self-healing patterns
- Implement monitoring and alerting as code (Terraform / GitOps) instead of manual UI configuration
- Reduce alert noise through event correlation, tagging standards and actionable alert design
- Build reusable data quality and observability frameworks, templates and accelerators adopted by multiple data product teams
- Lead post-incident reviews and translate learnings into systemic platform improvements
- Contribute to presales activities, client-facing assessments and internal training materials for the practice
Requirements
- 4+ years in SRE, DevOps, Data Engineering or Data Platform Operations, with at least 1 - 2 years focused on data platforms or data pipelines
- Solid SRE fundamentals: Golden Signals, SLI/SLO definition and calculation, error budgets and burn rate, incident lifecycle, ITIL basics
- Strong SQL and practical Python for automation and validation scripting
- Hands-on production experience with at least one cloud platform: Azure, AWS or GCP
- Experience with at least one modern data platform and orchestration stack - Databricks (preferred), Snowflake, Spark, Airflow, Azure Data Factory, dbt
- Working knowledge of observability tooling: New Relic, Datadog, Splunk, Elastic Stack, Grafana / OpenTelemetry (any two or more)
- Infrastructure as Code with Terraform (Ansible is a plus) and CI/CD experience (Azure DevOps, GitLab CI, GitHub Actions)
- Experience with incident and change management tooling: ServiceNow, PagerDuty or equivalent
- Ability to troubleshoot complex distributed data issues under SLA pressure
- B2+ English - the role is client-facing and requires clear written and spoken technical communication
Nice to have
- Databricks certification (Data Engineer Associate / Professional) or equivalent cloud data certification
- Hands-on experience with dedicated data observability platforms (Monte Carlo, Soda, Anomalo, Great Expectations)
- Data catalog and lineage tooling (Unity Catalog, OpenLineage, Collibra)
- FinOps: cloud, platform and telemetry cost optimization
- Exposure to AI/ML pipeline monitoring or AI Reliability Engineering
- Power BI or another BI layer, from a monitoring and reliability perspective
- Event correlation / AIOps experience (New Relic Decisions, IBM NOI, ServiceNow ITOM)
- Mentoring or team lead experience