- Work model
- Remote
- Experience
- 5+ years
- Employment
- Not specified
- Compensation
- Not disclosed
- Technology signal
- 12 tags
Technology context
12Parsed from the vacancy text; ordered by relevance to this role.
Full listing
Role description
Seeking Senior Data Engineers with deep experience in distributed data systems, Spark-based data processing, and production-grade data platform engineering. These engineers will be expected to operate independently, own technical solutions end-to-end, and contribute to both system design and implementation in a highly automated development environment.
What project we have for you
Client's Customer Platform & Data (CPD) organization is expanding its engineering teams to support critical initiatives across customer data, identity resolution, bookings, loyalty programs, and AI-powered customer insights. The platform processes large-scale batch and real-time data, powering customer-facing experiences and data-driven decision-making across Expedia Group.
Engineers will build, optimize, and operate production-grade data pipelines and platforms using Scala, Spark, Kafka, and Airflow. The environment emphasizes strong engineering discipline, data quality, system reliability, and AI-assisted software development practices.
What you will do
- Design, build, and optimize batch and streaming data pipelines using Scala and Spark.
- Develop and support Kafka/Flink-based streaming solutions.
- Build and maintain Airflow DAGs, backfills, and production workflows.
- Design data models, schemas, and source-to-target mappings.
- Implement data quality controls, validation, and monitoring.
- Troubleshoot production issues and ensure platform reliability.
- Review and validate AI-generated code and maintain engineering quality standards.
- Own systems end-to-end, including performance, cost, scalability, and reliability.
What you need for this
- 5+ years of Data Engineering experience.
- Strong Scala and Apache Spark.
- Experience building and owning production ETL/ELT pipelines.
- Streaming experience with Kafka, Kafka Streams, Flink, or similar.
- Apache Airflow.
- Comfortable with Java, Scala, Python, and configuration-heavy code.
- Strong data modeling, schema design, and source-to-target mapping.
- Data quality, validation, and production troubleshooting.
- Strong software engineering practices (testing, CI/CD, versioning).
- Ability to review AI-generated code.
Strong nice-to-have
- Flink expertise.
- ScyllaDB, Cassandra, DynamoDB, or other NoSQL platforms.
- Customer Data Platform (CDP), identity resolution, loyalty, clickstream, or booking data experience.
- Experience with SLAs, SLOs, observability, and monitoring.
- GitHub Copilot, Claude, Cursor, or similar AI-assisted development tools
Success Profile
The ideal candidate is a senior, highly autonomous engineer who can clearly speak:
- A large-scale Spark job they built or optimized
- Scala production experience
- Streaming experience with Kafka Streams, Flink, or similar
- Airflow DAGs, backfills, and production workflows
- Data quality checks and pipeline validation
- Comfortable with Java, Scala, Python, and configuration-heavy code.
- Production troubleshooting and deployment discipline
- Experience reviewing AI-generated code rather than blindly accepting it
- Ownership mindset over systems, datasets, reliability, cost, and quality
- The team operates spec-to-code methodology. So awareness of frameworks and tools like spec-led development, spec-kit, agent-skills etc. are great.