About the role
Runpod rents GPU capacity to developers running AI workloads and says the platform has processed more than 20 billion inference requests. This role owns LLM serving performance end to end, against the stated goal of making Runpod the fastest and most cost efficient place in the world to run inference.
It is measurement before optimisation, which is the right order. Defining how inference performance is measured, meaning throughput, time to first token, inter-token latency and cost per token, and building tooling that makes those numbers rigorous and repeatable. Then profiling the serving stack from scheduling and memory management down to kernels and interconnect, for large models on single-node and multi-node GPU deployments. What you learn has to come back as production runtimes, configurations and defaults that customers benefit from without asking.
Runpod closed a $100M Series A in June 2026 and describes itself as a small, remote-first team. Remote within the United States.
Skills
How to apply
Apply directly on the employer's application page. Your application goes straight to them.
More from Runpod
Jobs like this
Republished listing
This opportunity was discovered on Runpod's public careers page and is republished here for discovery purposes. Applications are handled by the employer.
If this is your role and you want it off the board, contact us.