- Work model
- Remote
- Experience
- 5+ years
- Employment
- Not specified
- Compensation
- Not disclosed
- Technology signal
- 15 tags
Technology context
15Parsed from the vacancy text; ordered by relevance to this role.
AIAI-Assisted DevelopmentPythonGCPCloudKafkaSparkSnowflakeDatabricksApache IcebergData Software EngineeringDatabricks Unity CatalogDelta LakeGoogle Cloud BigQuerySnowflake Horizon Catalog
Full listing
Role description
We are building a governed lakehouse platform with UniForm, dual-format pipelines, and zero-copy sharing across Snowflake and Databricks. As a Lead Data Software Engineer , you will define reusable integration patterns for BigQuery, Kafka/CDC, and catalogs while applying AI-assisted development tools across delivery.
Responsibilities
- Design and deliver a lakehouse UniForm write layer using dual-format metadata (Delta + Iceberg) that all target consumers can read without conversion
- Define and validate GCS-to-BigQuery ingestion pipeline patterns for structured operational domains including sales, delivery, schedule, and performance data
- Implement CDC patterns with Kafka to enable real-time and near-real-time data movement into the lakehouse
- Develop dependency-aware bookkeeping approaches and data lineage tracking patterns for consistent use across all data pipelines
- Ensure adapter code is modular, version-controlled, and built for reuse across new data source integrations
- Configure Iceberg external table definitions inside Snowflake's Horizon catalog
- Validate zero-copy read access from Snowflake to managed Iceberg / Delta tables without any data movement
- Implement and test tenant-scoped access controls aligned with Snowflake Horizon catalog metadata governance
- Implement and certify a Delta Sharing adapter that enables live, zero-copy sharing from Delta Lake tables to Databricks consumers
- Configure Delta Sharing endpoint registration and manage sharing agreement workflows
- Register Snowflake and Databricks as named connector types within the connector registry
- Implement RBAC, tenant-scoped authorization, and metering hooks that align with the billing framework for governed data-out flows
Requirements
- Proven experience of 5+ years in data engineering or software engineering focused on large-scale data platforms
- Deep expertise with GCP BigQuery, Apache Iceberg, and Delta Lake
- Hands-on proficiency in Python and Spark to develop and support data pipelines
- Solid understanding of data lake architecture, Iceberg UniForm, and Delta Sharing
- Working knowledge of Kafka/CDC patterns for real-time data movement
- Background integrating Snowflake Horizon catalog with Databricks Unity Catalog
- Active, practical experience with AI-assisted development tools such as Claude Code, GitHub Copilot, or Cursor
- Capability to demonstrate effective AI tooling usage during a technical screening
- English proficiency at B2 level or higher