About the role
Make open models fast and cheap to run on a network of distributed compute providers.
What you'll do
- Optimize LLM serving runtimes
- Benchmark models across GPU types
- Build quantization and batching pipelines
What we're looking for
- 3+ years in ML engineering
- Experience with vLLM, TensorRT or similar
- Strong Python
Nice to have
- CUDA experience
- Research background
Tech stack
PythonPyTorchvLLMCUDAKubernetes
Compensation
$70k – $130k + equity
Benefits
Meaningful equity
Fully remote
GPU credits
