Staff Software Engineer: AI Inference Data Plane

Digital Ocean $167.2K - $209K/year Posted about 3 hours ago

100% remote US only · US (remote)

About the role

DigitalOcean is expanding its AI Infrastructure layer to support the next generation of AI-driven applications. As a key technical leader on the AI Inference Data Plane team, you will design, develop, and deliver high-scale, resilient data plane services that power the company's Inference as a Service offering. You will work at the intersection of distributed systems and specialized AI hardware to ensure customers can deploy and scale their models with industry-leading performance and reliability. This is a hands-on role, developing high quality software while taking full advantage of the latest AI coding agents.

Key Responsibilities:

  • Act as a technical leader on the team, driving the end-to-end design, development, and delivery of critical data plane components hosting large generative AI models.
  • Architect and refine system design proposals for a high-scale, multi-tenant AI inference cloud ecosystem meeting rigorous availability and resiliency standards.
  • Implement and optimize distributed inference hosting using techniques like tensor/data parallelism, KV cache optimizations, and smart routing.
  • Build on Kubernetes-native distributed inference frameworks like llm-d (or alternatives such as NVIDIA Dynamo, Ray Serve, KServe) to deliver prefill/decode disaggregation, KV-cache-aware routing, tiered prefix caching, and wide expert parallelism for MoE models.
  • Solve the distributed-systems problems unique to LLM serving: inference-aware load balancing on queue depth, cache locality, and predicted latency; flow control and fairness across tenants; autoscaling inference pools; and moving gigabytes of KV-cache between prefill and decode instances with negligible overhead.
  • Contribute upstream to llm-d, vLLM, and the inference gateway ecosystem, representing DigitalOcean in these communities.
  • Coach and mentor junior engineers, and maintain critical high-scale services with observability tooling and well-defined SLOs.

Qualifications:

  • Hands-on experience hosting large language or multimodal models using inference engines like vLLM, SGLang, or TensorRT, including internals like continuous batching, paged attention, and prefix caching.
  • Familiarity with distributed inference serving frameworks such as llm-d, NVIDIA Dynamo, or Ray Serve.
  • Understanding of why cluster-scale serving is hard: partitioned KV-cache locality, routing that preserves cache hit rates and tail latency, and disaggregated prefill/decode with fast cross-pod KV transfer.
  • Knowledge of common LLM architectures and optimization techniques such as continuous batching and quantization.
  • Expert-level proficiency in Go or Python and familiarity with gRPC.
  • Proven experience shipping customer-facing software and running critical services in a high-scale environment, with merged contributions to vLLM, llm-d, SGLang, or similar projects strongly preferred.

Skills

How to apply

Apply directly on the employer's application page. Your application goes straight to them.

Republished listing

This opportunity was discovered on Digital Ocean's public careers page and is republished here for discovery purposes. Applications are handled by the employer.

Digital Ocean team? Claim this listing or ask us to remove it.