Vacancy catalog
EPAM
Open role>14 days

Senior Data Software Engineer with Databricks and Azure

EPAMKazakhstan
Work model
Remote
Experience
3+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
12 tags

Technology context

12

Parsed from the vacancy text; ordered by relevance to this role.

Full listing

Role description

We are looking for a Senior Data Software Engineer to join the Data Engineering CoE, building and scaling data pipelines in Databricks with PySpark and Python to deliver data products for the client's stakeholders.

Responsibilities

  • Support a team of data engineers to build pipelines used by MLOps and ML Engineers on ML modeling teams
  • Develop, optimize, and maintain data transformation pipelines in Databricks (PySpark)
  • Work with data stored in ADLS Gen2 and SAP HANA Data Lake, primarily in Delta/Parquet format
  • Implement and maintain data quality checks, including schema validation, deduplication, enrichment, and tagging
  • Communicate with stakeholders to understand business processes and model input data
  • Tune performance for large-scale datasets

Requirements

  • 3+ years of experience in data engineering or a related field
  • Proficiency in Python, PySpark, and Databricks, including Delta Lake
  • Familiarity with software version control tools such as GitHub and Git
  • Experience with CI/CD frameworks such as GitHub Actions
  • Knowledge of data lake technologies
  • Experience working with MS Azure
  • English proficiency at B2 level or higher

Nice to have

  • Familiarity with at least one other programming or scripting language, such as Java, SQL, or Scala