About the role
Runway builds world models, and this role owns model evaluation across the company, end to end.
The argument in the posting is the reason to read it. Every decision about a model, which checkpoint to keep training, what to ship to millions of users, which datamix and architecture look most promising, rests on evals, and that work is currently spread across the organisation. The job is to turn it into one platform: generate samples at scale, score them with automated metrics and human annotation, track results across checkpoints and releases, and put answers in front of researchers in minutes rather than days. It also includes defining what better means operationally, how confident you are in it, and how a result becomes a ship or no-ship decision.
The stack is Python and PyTorch, cloud native across multiple clouds, with Kubernetes and home-grown orchestration, and Runway says it is a heavy user of agentic development. The role is embedded in research teams as part of ML Platform. The band is the figure Runway publishes for US candidates. Remote.
Skills
How to apply
Apply directly on the employer's application page. Your application goes straight to them.
More from Runway
Jobs like this
Republished listing
This opportunity was discovered on Runway's public careers page and is republished here for discovery purposes. Applications are handled by the employer.
If this is your role and you want it off the board, contact us.