Senior Director, Inference Products and Optimizations
Digital Ocean $274.4K - $343K/year Posted 4 days ago
About the role
DigitalOcean's Inference Engine organization is seeking an experienced Senior Director of Engineering to lead a high-performing team building and scaling its Large Language Model inference products across the control plane, model optimization, and model architecture layers. You will own the inference product suite, spanning Serverless Inference, Dedicated Inference, Inference Router, Batch Inference, and Multimodal Inference, along with the model optimization and architecture stack that underpins them, delivering robust, cost-efficient systems that serve millions of users globally.
Key Responsibilities:
- Recruit, mentor, and coach engineers on the team, fostering a culture of ownership, technical excellence, and continuous improvement.
- Work with Product teams to define and execute the roadmap for all of DigitalOcean's inference products.
- Lead the design and evolution of the inference serving stack, driving deep technical strategy across vLLM, SGLang, and LLM-D to optimize throughput, latency, and GPU utilization at scale.
- Architect the model-serving and optimization layer, spanning quantization, KV-cache management, speculative decoding, and disaggregated serving, to deliver best-in-class performance per dollar.
- Collaborate with Product Management, other engineering teams, and key stakeholders to align priorities, manage dependencies, and communicate progress and risks.
- Ensure production health, stability, and on-call rotations that maintain customer SLAs.
- Institutionalize benchmarking frameworks, observability, and auto-tuning capabilities, and encourage contributions to open-source inference engines.
Qualifications:
- 10+ years of software engineering experience, with 6+ years in a technical leadership or management role, ideally within inference systems or AI/ML systems.
- Deep expertise in distributed systems design, Kubernetes at scale, LLM inference, and AI workload orchestration, scheduling, and resource management.
- Ability to engage in deep technical discussions on scalable control plane design, inference engines (vLLM, SGLang), and model architectures.
- Strategic knowledge of GPU architectures (NVIDIA and/or AMD), interconnects like NVLink, and hardware topology and their impact on AI training and inference performance.
- Expertise in defining and operationalizing deep inference metrics (TTFT, TPOT) to drive performance improvements and meet SLOs.
- Demonstrated ability to translate complex technical requirements into user-focused product features, plus excellent communication skills to align diverse teams around a shared vision.
Skills
How to apply
Apply directly on the employer's application page. Your application goes straight to them.
Republished listing
This opportunity was discovered on Digital Ocean's public careers page and is republished here for discovery purposes. Applications are handled by the employer.
Digital Ocean team? Claim this listing or ask us to remove it.