Open roleNew
Senior Data Software Engineer with AWS and Terraform
EPAMGeorgia; Armenia; Kazakhstan; Kyrgyzstan; Uzbekistan
- Work model
- Remote
- Experience
- 3+ years
- Employment
- Contract
- Compensation
- Not disclosed
- Technology signal
- 13 tags
Technology context
13Parsed from the vacancy text; ordered by relevance to this role.
AILLMAWSGenAIPythonAPITerraformSnowflakeCI/CDAutomationAirflowOOPData Software Engineering
Full listing
Role description
We are seeking a Senior Data Software Engineer to join a client-facing delivery team building and hardening cloud-native data pipelines on AWS as part of a data platform modernization program. The role involves ingesting and transforming large datasets with PySpark on AWS Glue and delivering curated, validated data into Snowflake, with a core focus on data quality, validation, and reconciliation for downstream analytics. This position is delivered at a Senior Consultant level with high autonomy and direct client stakeholder communication.
Responsibilities
- Design, build, and optimize scalable batch and incremental ETL/ELT pipelines using PySpark on AWS Glue
- Configure Glue jobs, crawlers, triggers, connections, bookmarks, workflows, and the Glue Data Catalog
- Tune workers, partitioning, and shuffle behavior for cost and performance optimization
- Model and load curated datasets into Snowflake with staging, transformation, and publishing layers
- Implement automated data quality and validation frameworks, including schema/contract enforcement and null/uniqueness/referential checks
- Develop row-count and financial reconciliation processes, anomaly detection, and quarantine/reject handling
- Configure and extend Glue Data Quality (DQDL) rules per requirements
- Write clean, modular, testable Python with unit/integration tests and reusable libraries
- Integrate pipelines with AWS services such as S3, IAM, Lambda, Athena, CloudWatch, Step Functions, and Secrets Manager
- Instrument observability through logging, metrics, alerting, and pipeline SLA monitoring
- Participate in code reviews, CI/CD automation, and documentation
- Engage directly with client stakeholders in requirements refinement, design walkthroughs, status reporting, and act as technical advisor within the workstream
Requirements
- 3+ years of experience with Python for production-level data engineering, including OOP and functional patterns
- Expertise in PySpark for distributed data processing and the DataFrame API
- Advanced proficiency in Snowflake, including data warehousing and staging/transformation layers
- Skills in AWS Glue, including job configuration, crawlers, Data Catalog, and DQDL
- Background in data quality engineering, including validation frameworks and reconciliation
- Proficiency in AWS services including S3, IAM, Lambda, Athena, and CloudWatch
- English proficiency at B2 level or higher
Nice to have
- Familiarity with Generative AI / LLM concepts
- Knowledge of Airflow / Step Functions orchestration
- Familiarity with Great Expectations or similar data quality frameworks
- Knowledge of Terraform / CloudFormation