Vacancy catalog
EPAM
Open role

Senior Infrastructure Engineer

EPAMArmenia; Georgia; Kazakhstan; Kyrgyzstan; Uzbekistan
Work model
Remote
Experience
5+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
21 tags

Technology context

21

Parsed from the vacancy text; ordered by relevance to this role.

AWSDistributed SystemsPythonKubernetesCloud NativeCloudGoReliability EngineeringCloud SecurityKafkaDevOpsRedisPostgreSQLTerraformLinuxSRECI/CDPrometheusGrafanaHTMLLoad Testing

Full listing

Role description

We are seeking a Senior Infrastructure Engineer with deep experience in cloud infrastructure, Linux internals, large-scale distributed systems, workload isolation, and a strong security mindset to join our team and help design, build, and operate a global, resilient platform.

Responsibilities

  • Design, build, and maintain a global, distributed, and resilient cloud infrastructure
  • Collaborate with infrastructure and product engineering teams to plan and deliver complex platform initiatives
  • Participate in architecture reviews, incident response, and performance analysis to ensure system reliability
  • Manage and provision AWS infrastructure using Terraform and Kubernetes
  • Write and maintain Kubernetes manifests and deployment configurations for critical workloads, including pod security contexts, anti-affinity rules, network policies, autoscaling, and health probes
  • Drive Production Readiness Reviews (PRR) for all new services, covering security, HA, performance, and observability gates
  • Design and operate multi-layer workload isolation using Linux kernel primitives: namespaces (pid, net, mnt, user, uts, ipc), cgroups, seccomp profiles, and capabilities, as the baseline security boundary
  • Evaluate and operate gVisor and Firecracker for workloads requiring hard tenant boundaries and near-native performance
  • Design and maintain Grafana dashboards
  • Manage Prometheus and VictoriaMetrics pipelines; define and tune P1/P2/P3 alert thresholds with runbooks
  • Contribute to and extend the internal k6-based load testing framework (load-testing-framework / library/k6/webhooks)
  • Design load scenarios using constant-arrival-rate profiles; instrument custom metrics (job_succeeded_count, job_failed_count) tagged by testid for Grafana correlation; stream test metrics to Prometheus via remote write; generate and publish HTML reports to file storage after each run

Requirements

  • 5+ years of experience in infrastructure, SRE, or platform engineering roles
  • Expertise in distributed systems and cloud-native architectures
  • Understanding of Linux internals: namespaces, cgroups, seccomp, capabilities, and system-level performance tuning
  • Experience operating infrastructure on AWS at scale
  • Proficiency in Terraform and Kubernetes, including security hardening of manifests
  • Experience designing and running load tests (k6, Gatling, Locust, or similar)
  • Understanding of network security and cloud security best practices
  • Excellent analytical, troubleshooting, and communication skills
  • Proficiency in English at a B2+ level

Nice to have

  • Skills in Golang/Python for building internal tooling
  • Familiarity with Kafka, Redis, ClickHouse, or PostgreSQL
  • Familiarity with observability tools such as Grafana, Prometheus, or VictoriaMetrics
  • Knowledge of encryption key hierarchies (CMK/DEK/KMS patterns) and HashiCorp Vault
  • Hands-on production experience with gVisor or Firecracker
  • Prior work extending or maintaining an internal testing framework
  • Contributions to or maintenance of open-source infrastructure projects