OLD
Long-running vacancy
This listing is older than 30 days but its source has not removed it. Verify availability on the original company page before applying.
Open role>1 month
GPU & ML Infrastructure Engineer
Svitla SystemsArgentina; Mexico; Chile; United States
- Work model
- Remote
- Experience
- Not specified
- Employment
- Full Time
- Compensation
- Not disclosed
- Technology signal
- 9 tags
Technology context
9Parsed from the vacancy text; ordered by relevance to this role.
Full listing
Role description
Svitla Systems Inc. is looking for a GPU & ML Infrastructure Engineer for a full-time position (40 hours per week) in the USA. Our client is a stealth startup. The successful candidate will own the end-to-end data generation process for GPU systems, including benchmark workloads, automated deployment, hardware telemetry collection, data quality validation, and dataset delivery. The role requires working with multiple NVIDIA data-center GPU generations and embedded or edge platforms.
- Strong experience deploying LLM inference and training workloads on GPUs, including quantized models.
- Experience diagnosing sensor, logging, and sampling issues in time-series hardware data.
- Ability to build reproducible GPU workloads and control sources of run-to-run variation.
- Strong Linux systems knowledge, including GPU driver stacks, process orchestration, scheduling, and timing.
- Experience collecting hardware telemetry programmatically using NVML, DCGM, BMC, IPMI, or Redfish.
- Strong Python skills for automation, telemetry collection, and data processing.
- Experience automating workload deployment and data collection across different hardware platforms.
- Ability to work independently and take ownership of technical processes.
- Availability to overlap with the client's team until 11:00 a.m. PST.
- Experience building data collection pipelines for hardware testing or systems research.
- Familiarity with GPU benchmarking, stress testing, and benchmark methodology.
- Knowledge of GPU power, thermal management, multi-GPU scaling, and NCCL.
- Experience building automated data quality checks for time-series or sensor data.
- Port existing test procedures to new data-center GPUs and edge devices.
- Build and maintain GPU benchmark workloads, including synthetic kernels and LLM inference and training.
- Ensure telemetry collection is accurate and consistent across platforms and data sources.
- Investigate sampling issues, timestamp inconsistencies, missing sensor data, and logging anomalies.
- Automate workload deployment, execution, data collection, and environment cleanup.
- Build automated checks for missing samples, irregular intervals, clock mismatches, and invalid telemetry.
- Deliver datasets in a consistent and documented format with complete run metadata.
- Document hardware targets, configurations, procedures, driver versions, and firmware versions.