Vacancy catalog
Intellias
Open role

Senior Data Scientist

IntelliasIntellias
Work model
Remote
Experience
5+ years
Employment
Not specified
Compensation
Not disclosed
Technology signal
11 tags

Technology context

11

Parsed from the vacancy text; ordered by relevance to this role.

Full listing

Role description

Let's breathe life into great tech ideas! With 3,000 people globally, Intellias is a company where benchmark technological solutions are born. Join in and take your part in digitalizing the world.

What project we have for you

Join a transformative data and AI platform initiative aimed at modernizing enterprise-scale capabilities and enabling real-time decision-making. This project delivers a comprehensive roadmap covering AI, MLOps, data governance, and platform scalability, supporting a shift towards data-first operations and intelligent automation.

What you will do

  • Platform setup. With customer's platform team, provision and configure the Databricks environment in customer's cloud account: catalog and permissions, compute (including GPU), secrets, experiment tracking.
  • Secure data ingestion. Build repeatable ingestion of raw verification outputs and document images: parsing, pseudonymisation, data-quality checks, schema handling.
  • Feature engineering at scale. Implement the feature tables:
  • batch image-embedding jobs;
  • exact and approximate linkage keys;
  • cross-transaction and velocity aggregates;
  • pseudonymous entity resolution.
  • Modelling support. Implement feature-selection and mixed-data clustering methods. Run model comparisons and stability tests, and support density-model training runs.
  • Automation. Orchestrate the end-to-end pipeline as a scheduled, idempotent, re-runnable workflow.
  • Monitoring. Set up the data-quality and drift-monitoring prototype, with alerting, on feature and score tables.
  • Privacy engineering. Implement access controls, provenance fields, lineage and the scripted deletion procedure. Confirm that no raw personal data reaches the analytic layer.
  • Documentation and handover. Write runbooks and give a live walkthrough, so engineers can operate everything independently

What you need for this

  • PySpark & Spark SQL: Nested-JSON flattening (explode, structs),`mapInPandas` / pandas UDFs, window functions for velocity and distinct-count aggregates, partitioning and performance tuning
  • Databricks: Unity Catalog (catalogs, schemas, grants, column masks, tags, lineage), external locations on S3, cluster / compute policies incl. GPU, Auto Loader, Delta (MERGE, time travel, VACUUM), Workflows, Repos, secret scopes
  • AWS: S3, IAM roles / instance profiles, KMS decryption in jobs, Secrets Manager, basic cost awareness
  • Data engineering for ML: Medallion design (Bronze / Silver / Gold), schema evolution, idempotent loads, DQ profiling, Feature Engineering in UC (feature tables)
  • Pseudonymisation in pipelines: HMAC tokenization with per-key-type keys, normalisation (Unicode NFKD, case, token sort), phonetic keys, never logging raw values, scripted deletion (DROP / VACUUM / secret destruction)
  • Entity resolution (implementation): Blocking / candidate generation with exact and approximate keys, edit-distance tolerance, match scores, surrogate IDs, conflict counting rather than merging
  • Image processing at scale: Running a GPU batch job that decrypts, crops and embeds images in memory (PyTorch / timm), PCA, LSH bucketing
  • Mixed-data clustering: Hands-on: `kmodes` (k-prototypes), `gower` + `kmedoids` with sampling, R `kamila` on Databricks, `StepMix` (latent class), bootstrap ARI
  • Feature selection: Mutual information, correlation pruning, PCA. Can run a BPSO / GA search with`mealpy`
  • MLOps & monitoring: M*Lflow experiments and registry (aliases), Lakehouse Monitoring * (profile and drift metrics on Delta tables), SQL alerts, PSI / JS divergence
  • Handover: Writes runbooks and README-level documentation. Leaves a single end-to-end Workflow that someone else can run

Nice to have

  • Experience with AssureID / identity-verification JSON outputs.
  • PyOD and basic PyTorch training (to support A on the VAEs in week 2).