All positions

ML Inference Engineer

AI / ML · Full-time · Remote · Mid-level

Posted Oct 7, 2026

About the role

Make open models fast and cheap to run on a network of distributed compute providers.

What you'll do

  • Optimize LLM serving runtimes
  • Benchmark models across GPU types
  • Build quantization and batching pipelines

What we're looking for

  • 3+ years in ML engineering
  • Experience with vLLM, TensorRT or similar
  • Strong Python

Nice to have

  • CUDA experience
  • Research background

Tech stack

PythonPyTorchvLLMCUDAKubernetes

Compensation

$70k – $130k + equity

Benefits

Meaningful equity
Fully remote
GPU credits

Think you'd be a great fit?