Vacancy catalog
EPAM
Open role

Lead Data Integration Engineer

EPAMArgentina; Chile; Colombia
Work model
Remote
Experience
5+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
20 tags

Technology context

20

Parsed from the vacancy text; ordered by relevance to this role.

Full listing

Role description

We are seeking a Lead Data Integration Engineer to deliver production dbt and Snowflake models powering GTM analytics and AI agents. You will create trusted semantic layers, unify Salesforce and Gong-style interaction data, and prepare datasets for Snowflake Cortex and Snowflake Intelligence.

Responsibilities

  • Design and maintain production data models in dbt and Snowflake to measure account health, engagement velocity, and customer intent signals
  • Integrate structured records with unstructured interaction data (Gong call transcripts, emails, Salesforce CRM objects, and product usage telemetry) into reliable account-level dimensions
  • Define consistent modeling patterns, naming conventions, and metric definitions across GTM datasets
  • Implement and operate an enterprise semantic layer (such as the dbt Semantic Layer or MetricFlow) to prevent conflicting metric calculations across business tools
  • Prepare clean, context-rich datasets optimized for retrieval, search, and reasoning in Snowflake Cortex, Snowflake Intelligence, and internal AI agents
  • Create feature tables and context stores that enable LLM agents for automated meeting prep, competitive intelligence, and account research
  • Administer metadata and data enrichment patterns so AI agents correctly interpret account stages, customer sentiment, and relationship history
  • Develop automated tests, schema validations, and freshness monitors in dbt for both structured metrics and unstructured data transformations
  • Enforce data governance, role-based access controls (RBAC), and audit trails across sensitive sales pipeline and customer interaction data
  • Collaborate with Revenue Operations and governance teams to document data lineage and maintain clear dataset ownership
  • Contribute within an agile sprint cadence, supporting backlog planning, design reviews, and technical standards
  • Deliver self-service data models and documentation so analysts and product teams can query data independently

Requirements

  • Proven track record with 5+ years of commercial experience in analytics engineering, data engineering, or data architecture roles
  • Expert-level proficiency in advanced SQL, including window functions, CTEs, complex joins, analytical queries, and execution plan optimization
  • Extensive hands-on experience with 5+ years of production dbt work, including modular project structure, Jinja macros, custom tests, incremental models, and snapshots
  • Solid practical experience with Snowflake, including virtual warehouse sizing, query profile analysis, clustering keys, Time Travel, zero-copy cloning, and access control
  • Demonstrated experience modeling B2B SaaS revenue and interaction data (Salesforce CRM, Gong, Marketo, or subscription billing data)
  • Deep understanding of Kimball dimensional modeling (star schemas, slowly changing dimensions SCD 1/2/4, conformed dimensions) and semantic layer design
  • Daily comfort with Git workflows, branching strategies, code reviews, and automated CI test runs
  • Willingness to use AI-assisted development tools (GitHub Copilot, Cursor, Claude Code) to accelerate modeling work, test coverage, and documentation
  • English proficiency at B2 (Upper-Intermediate) level or higher

Nice to have

  • Working familiarity with Snowflake Cortex AI functions (Cortex Search, LLM functions, embeddings)
  • Understanding of vector embeddings, chunking strategies, and retrieval-augmented generation (RAG) for conversational text
  • Working knowledge of Python for data manipulation (Pandas), API extraction, or pipeline automation
  • Experience with metadata and lineage tools such as Alation, Collibra, DataHub, or dbt Docs
  • Background designing self-service data models for Tableau, Looker, or Power BI