Vacancy catalog
EPAM
Open roleNew

Principal - AI & HPC Data Centre Compute

EPAMUK, London
Work model
Hybrid
Experience
12+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
6 tags

Technology context

6

Parsed from the vacancy text; ordered by relevance to this role.

Node.jsAIKubernetesPyTorchInfiniBandAutomation

Full listing

Role description

We're looking for a Director - AI & HPC Data Centre Compute to join our team in London, United Kingdom in a hybrid working mode.

In this senior leadership role, you will drive EPAM's AI and HPC data centre strategy, leading engagements that optimise compute infrastructure across full-stack environments - from accelerators and networking to schedulers, containers and AI platforms. You will address cost, scalability and power constraints through software engineering, platform optimisation and systems integration rather than hardware procurement, helping clients deliver performance and efficiency at scale.

Responsibilities

  • Advise hyperscalers, neoclouds and enterprises on compute strategies for AI and HPC, including power, cooling and network design
  • Optimise AI training and inference workloads (LLMs, multimodal and scientific AI) across distributed clusters for cost, throughput and latency targets
  • Lead HPC cluster design and orchestration using Slurm, Kubernetes and parallel processing models (MPI)
  • Enhance GPU utilisation by addressing bottlenecks across compute, memory and data pipelines; collaborate with energy teams on power-aware scheduling
  • Architect scalable AI platforms integrating MLOps frameworks, automation tools and reusable assets
  • Develop repeatable AI/HPC offerings, contribute to solution roadmaps and establish strategic partnerships with major ecosystem players
  • Track industry trends in AI factories, GPU economics and liquid cooling technologies, representing EPAM at industry events and thought leadership forums

Requirements

  • 12+ years of experience in HPC, AI infrastructure, accelerated computing or distributed systems architecture
  • Deep knowledge of GPU architectures, AI workloads, networking and large-scale cluster operations
  • Hands-on expertise with Slurm, Kubernetes and performance optimisation across multi-node environments
  • Proven ability to advise senior stakeholders and influence technical strategy at C-level
  • Demonstrated experience in pre-sales solutioning and shaping complex technology engagements

Nice to have

  • Exposure to NVIDIA ecosystems (DGX, HGX, SuperPOD) or alternative accelerators (AMD or similar)
  • Familiarity with InfiniBand/RoCE networking, PyTorch, high-density rack design, liquid cooling or GPU-as-a-Service deployments