Vacancy catalog
EPAM
Open role

Senior Real-Time Observability Engineer

EPAMRomania
Work model
Remote
Experience
5+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
14 tags

Technology context

14

Parsed from the vacancy text; ordered by relevance to this role.

TypeScriptDistributed SystemsJavaPythonKubernetesAlgorithmsCloud NativeCloudC++APIRustCI/CDAdvanced Systems EngineeringOpenTelemetry

Full listing

Role description

We are seeking a Senior Real-Time Observability Engineer to build and evolve the systems that provide unified visibility into a distributed, latency-sensitive Real-Time platform, working across telemetry pipelines, synthetic monitoring, and GitOps-driven observability infrastructure.

Responsibilities

  • Build and maintain application components that aggregate, correlate, and present observability data from individual services into unified, end-to-end views of the Real-Time platform
  • Implement customer-centric aggregation and hotspot detection algorithms that reflect timeliness, completeness, accuracy, and stability of market data across the Real-Time estate
  • Develop and extend telemetry pipelines that ingest, transform, and route metrics, traces, and logs from distributed services into a coherent observability layer
  • Design and build custom synthetic monitoring agents deployed across hundreds of global sites to continuously measure customer experience of Real-Time data from the edge
  • Implement GitOps/API-driven workflows for observability assets to ensure consistent deployment, versioning, and promotion through build pipelines
  • Collaborate with squads to integrate observability instrumentation into applications throughout Dev, Test, PPE, and Prod environments

Requirements

  • 5+ years of hands-on software engineering experience building data-pipeline applications, including metrics aggregation, distributed tracing, or real-time streaming
  • Background in latency-sensitive or market data systems
  • Proficiency in at least one systems-level language (C++, Go, or Rust) and one scripting/application language (Python, Java, or TypeScript)
  • Experience in C++ to understand existing applications
  • English proficiency at B2 level or higher

Nice to have

  • Experience building and operating custom synthetic monitoring solutions at scale, including lightweight agents distributed across geographically diverse sites
  • Familiarity with OpenTelemetry SDKs, collectors, and schema conventions for instrumentation and telemetry export
  • Skills in cloud-native, containerized, and Kubernetes environments, including deploying and operating services at scale
  • Proficiency with API- and GitOps-based workflows, config-as-code, and CI/CD pipelines for infrastructure
  • Strong analytical mindset for modeling complex distributed-system behaviors and understanding customer impact, along with effective communication skills to work across squads and simplify system performance concepts for broader stakeholders