Vacancy catalog
EPAM
Open roleNew

Senior Data Software Engineer with AWS and Terraform

EPAMGeorgia; Armenia; Kazakhstan; Kyrgyzstan; Uzbekistan
Work model
Remote
Experience
3+ years
Employment
Contract
Compensation
Not disclosed
Technology signal
13 tags

Technology context

13

Parsed from the vacancy text; ordered by relevance to this role.

AILLMAWSGenAIPythonAPITerraformSnowflakeCI/CDAutomationAirflowOOPData Software Engineering

Full listing

Role description

We are seeking a Senior Data Software Engineer to join a client-facing delivery team building and hardening cloud-native data pipelines on AWS as part of a data platform modernization program. The role involves ingesting and transforming large datasets with PySpark on AWS Glue and delivering curated, validated data into Snowflake, with a core focus on data quality, validation, and reconciliation for downstream analytics. This position is delivered at a Senior Consultant level with high autonomy and direct client stakeholder communication.

Responsibilities

  • Design, build, and optimize scalable batch and incremental ETL/ELT pipelines using PySpark on AWS Glue
  • Configure Glue jobs, crawlers, triggers, connections, bookmarks, workflows, and the Glue Data Catalog
  • Tune workers, partitioning, and shuffle behavior for cost and performance optimization
  • Model and load curated datasets into Snowflake with staging, transformation, and publishing layers
  • Implement automated data quality and validation frameworks, including schema/contract enforcement and null/uniqueness/referential checks
  • Develop row-count and financial reconciliation processes, anomaly detection, and quarantine/reject handling
  • Configure and extend Glue Data Quality (DQDL) rules per requirements
  • Write clean, modular, testable Python with unit/integration tests and reusable libraries
  • Integrate pipelines with AWS services such as S3, IAM, Lambda, Athena, CloudWatch, Step Functions, and Secrets Manager
  • Instrument observability through logging, metrics, alerting, and pipeline SLA monitoring
  • Participate in code reviews, CI/CD automation, and documentation
  • Engage directly with client stakeholders in requirements refinement, design walkthroughs, status reporting, and act as technical advisor within the workstream

Requirements

  • 3+ years of experience with Python for production-level data engineering, including OOP and functional patterns
  • Expertise in PySpark for distributed data processing and the DataFrame API
  • Advanced proficiency in Snowflake, including data warehousing and staging/transformation layers
  • Skills in AWS Glue, including job configuration, crawlers, Data Catalog, and DQDL
  • Background in data quality engineering, including validation frameworks and reconciliation
  • Proficiency in AWS services including S3, IAM, Lambda, Athena, and CloudWatch
  • English proficiency at B2 level or higher

Nice to have

  • Familiarity with Generative AI / LLM concepts
  • Knowledge of Airflow / Step Functions orchestration
  • Familiarity with Great Expectations or similar data quality frameworks
  • Knowledge of Terraform / CloudFormation