- Work model
- Remote
- Experience
- 5+ years
- Employment
- Not specified
- Compensation
- Not disclosed
- Technology signal
- 9 tags
Technology context
9Parsed from the vacancy text; ordered by relevance to this role.
Full listing
Role description
We are looking for a Senior/Lead AI/ML Performance Engineer to benchmark and profile AI/ML workloads on TPU hardware to generate the empirical performance data that feeds the optimization solver's cost model.
Responsibilities
- Conduct JAX/XLA benchmarking on physical TPU slices
- Capture and analyze TPU Profiler traces
- Build hardware coefficient matrices for use in the optimizer's cost model
- Profile LLM training/inference performance (FLOPS, memory access, token throughput)
- Collaborate with Optimization Engineers to calibrate solver inputs
Requirements
- Strong understanding of the modern AI/ML landscape and LLM architectures
- Hands-on experience with LLM models/pipelines (training and inference)
- Proficiency in Python and/or Golang
- Familiarity with TensorFlow, PyTorch, and LangChain
- Experience with JAX/XLA and TPU-specific profiling tools
Nice to have
- Experience with GPU/TPU performance benchmarking at scale
- Background in ML systems or MLPerf-style benchmarking