Vacancy catalog
EPAM
Open roleNew

Platform Engineering Team Lead

EPAMIsrael
Work model
Hybrid
Experience
3+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
20 tags

Technology context

20

Parsed from the vacancy text; ordered by relevance to this role.

AIAWSMachine LearningAzurePythonGCPKubernetesData ScienceCloudKafkaDockerSparkTerraformLinuxDatabricksRESTAnsibleAutomationSQLPlatform Engineering

Full listing

Role description

We are seeking an experienced and visionary Platform Engineering Team Lead to lead a team of highly skilled platform experts. In this role, you will own the architecture, stability, and scaling of our enterprise data streaming and processing platforms.

You will act as a key technical leader, bridging the gap between Infrastructure, Data Engineering, and Data Science, ensuring high availability, continuous automation, and state-of-the-art platform observability.

Responsibilities

  • Team Leadership: Provide professional and personal management, mentorship, and guidance to a team of platform engineering experts
  • Confluent/Kafka Platform Ownership: End-to-end responsibility for the Confluent Suite, including Apache Kafka, Kafka Connect, Schema Registry, REST Proxy, Confluent Cloud, and KSQL
  • Data Platforms: Own and manage enterprise data analytics platforms, including Azure Synapse
  • DataOps, MLOps, & Compute: Manage end-to-end infrastructure supporting DataOps, MLOps, and specialized GPU compute resources
  • Ingestion & Logging Infrastructure: Oversee streaming ingestion and logging pipelines based on Apache NiFi, PortX, Fluent Bit, and Filebeat
  • Lifecycle & Projects: Lead complex architecture, installation, upgrading, and migration projects for all data platforms
  • Incident Management: Drive deep troubleshooting and resolution of complex production issues
  • Architectural Guidance: Provide professional architectural advisory and support to Data Engineering and Data Science development teams
  • Vendor Management: Interface directly with external vendors and technology partners
  • Automation & Observability: Drive automation initiatives and implement advanced system monitoring and observability frameworks
  • System Resiliency: Lead initiatives for Capacity Planning, High Availability (HA), and Disaster Recovery (DR)

Requirements

  • Leadership: At least 3 years of experience leading and managing a technology/engineering team
  • Domain Expertise: At least 5 years of experience in provisioning, maintaining, and operating large-scale Data and/or Streaming platforms
  • Operating Systems: Deep, hands-on experience working in Linux environments
  • Kafka & Confluent: Significant hands-on experience with on-premises Apache Kafka / Confluent and all its core ecosystem components
  • Production Operations: Proven track record in troubleshooting and root-cause analysis of complex issues in high-pressure Production environments
  • Cross-Functional Collaboration: Extensive experience working closely with Software/Data Developers and Solutions Architects
  • Big Data Ecosystem: Strong hands-on experience with at least one or more of the following: Hadoop, Cloudera, Databricks, Apache Spark, or Azure Synapse
  • Configuration Management: Solid experience working with Ansible for automation
  • Modern Paradigms: Hands-on experience in at least one of these domains: DataOps, MLOps, or AI/ML Platforms
  • A systemic, holistic approach to system architecture and planning
  • Excellent self-learning capabilities and adaptability to new technologies
  • Outstanding interpersonal skills with a strong service-oriented mindset
  • Proven ability to work effectively across multiple cross-functional departments (interfaces)

Nice to have

  • Data Ingestion: Strong hands-on experience with Apache NiFi (Highly Advantageous)
  • Cloud Infrastructure: Experience working within enterprise-grade Cloud environments (AWS, Azure, or GCP)
  • Distributed Querying: Experience with Trino (Presto SQL)
  • Infrastructure as Code (IaC): Experience with Terraform
  • Containerization: Experience with containers and container orchestration (Docker, Kubernetes)
  • Managed Streaming: Practical experience with Confluent Cloud
  • Scripting & Big Data Development: Practical experience with Python and PySpark
  • Scale: Prior experience working within a large Enterprise organization