Vacancy catalog
OLD

Long-running vacancy

This listing is older than 30 days but its source has not removed it. Verify availability on the original company page before applying.

Svitla Systems
Open role>1 month

GPU & ML Infrastructure Engineer

Svitla SystemsArgentina; Mexico; Chile; United States
Work model
Remote
Experience
Not specified
Employment
Full Time
Compensation
Not disclosed
Technology signal
9 tags

Technology context

9

Parsed from the vacancy text; ordered by relevance to this role.

Full listing

Role description

Svitla Systems Inc. is looking for a GPU & ML Infrastructure Engineer for a full-time position (40 hours per week) in the USA. Our client is a stealth startup. The successful candidate will own the end-to-end data generation process for GPU systems, including benchmark workloads, automated deployment, hardware telemetry collection, data quality validation, and dataset delivery. The role requires working with multiple NVIDIA data-center GPU generations and embedded or edge platforms.

  • Strong experience deploying LLM inference and training workloads on GPUs, including quantized models.
  • Experience diagnosing sensor, logging, and sampling issues in time-series hardware data.
  • Ability to build reproducible GPU workloads and control sources of run-to-run variation.
  • Strong Linux systems knowledge, including GPU driver stacks, process orchestration, scheduling, and timing.
  • Experience collecting hardware telemetry programmatically using NVML, DCGM, BMC, IPMI, or Redfish.
  • Strong Python skills for automation, telemetry collection, and data processing.
  • Experience automating workload deployment and data collection across different hardware platforms.
  • Ability to work independently and take ownership of technical processes.
  • Availability to overlap with the client's team until 11:00 a.m. PST.
  • Experience building data collection pipelines for hardware testing or systems research.
  • Familiarity with GPU benchmarking, stress testing, and benchmark methodology.
  • Knowledge of GPU power, thermal management, multi-GPU scaling, and NCCL.
  • Experience building automated data quality checks for time-series or sensor data.
  • Port existing test procedures to new data-center GPUs and edge devices.
  • Build and maintain GPU benchmark workloads, including synthetic kernels and LLM inference and training.
  • Ensure telemetry collection is accurate and consistent across platforms and data sources.
  • Investigate sampling issues, timestamp inconsistencies, missing sensor data, and logging anomalies.
  • Automate workload deployment, execution, data collection, and environment cleanup.
  • Build automated checks for missing samples, irregular intervals, clock mismatches, and invalid telemetry.
  • Deliver datasets in a consistent and documented format with complete run metadata.
  • Document hardware targets, configurations, procedures, driver versions, and firmware versions.