Vacancy catalog
EPAM
Open role

Data Software Engineer

EPAMMexico, Mexico City
Work model
Hybrid
Experience
2+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
10 tags

Technology context

10

Parsed from the vacancy text; ordered by relevance to this role.

AIAWSPythonSparkDatabricksSQLAmazon S3AWS GlueAWS LambdaData Software Engineering

Full listing

Role description

We are seeking a Data Software Engineer to join our team. In this role, you will contribute to building modern data solutions that support scalable analytics and data-driven decision-making. You will collaborate with cross-functional teams to deliver reliable and efficient data infrastructure aligned with business needs.

Responsibilities

  • Design and implement robust data pipelines to support various analytical and operational use cases
  • Develop scalable and optimized ETL/ELT processes within the DATIO and AWS ecosystems
  • Integrate structured and unstructured data sources to enable unified data access and analysis
  • Implement batch and real-time processing strategies to handle diverse data workloads

Requirements

  • At least 2 years of relevant professional experience in data engineering or a similar area
  • Hands-on experience with DATIO - ADA for developing data solutions within enterprise environments
  • Strong Python skills, including managing libraries such as Pandas and PySpark for data processing and analytics
  • Proficiency in SQL, with expertise in query optimization and data modeling
  • Desirable experience in data processing using Spark for handling large-scale datasets
  • Working knowledge of AWS services, including Apache Spark and Hadoop for distributed data processing
  • Excellent written and verbal command of English and Spanish at a B2 level or higher

Nice to have

  • Experience with AWS S3 for storing both structured and unstructured data at scale
  • Familiarity with AWS Glue for building serverless ETL processes and managing data catalogs
  • Practical exposure to AWS Lambda for implementing event-driven and serverless data workflows