Staff Production Engineer
GitHub $140.4K - $372.3K/year Posted 3 days ago
About the role
GitHub is building a new Production Engineering organization to raise the operational bar across the platform. Production Engineers are software engineers who embed with product and infrastructure teams to improve reliability, performance, scalability, and developer productivity through software, automation, and operational excellence. As software development becomes increasingly AI-assisted, the reliability and operability of GitHub's platform have never been more important.
Responsibilities
- Embed with engineering teams to improve the reliability, scalability, and operability of GitHub's production systems.
- Own complex production problems end-to-end, from debugging live incidents to designing long-term engineering solutions.
- Design, build, and operate software and infrastructure that improves how GitHub runs in production.
- Partner with product and infrastructure teams to improve observability, capacity planning, performance, resiliency, and operational readiness.
- Write and review production-quality code, automate operational workflows, and eliminate repetitive manual work.
- Participate in incident response and drive learning through incident reviews and follow-up engineering investments.
- Influence technical direction across teams by establishing engineering practices that improve operational excellence at scale.
- Mentor engineers and help build Production Engineering as a technical discipline within GitHub.
Qualifications
- 9+ years of experience in software engineering or a related technical discipline delivering production software, or an equivalent combination of education and experience.
- Experience designing, operating, and debugging large-scale distributed systems, with a proven ability to drive improvements in high availability, automation, observability, and overall system reliability.
- Strong software engineering experience in one or more modern programming languages such as Go, Rust, C++, Java, Python, or C#.
- Experience with Linux, networking fundamentals, and systems performance.
- Demonstrated ability to influence technical direction across teams without direct authority.
- Preferred: experience operating services at scale, with Kubernetes, cloud infrastructure, service networking, databases, or distributed storage systems.
- Preferred: experience leading incident response and driving operational improvements across organizations.
Skills
How to apply
Apply directly on the employer's application page. Your application goes straight to them.
Republished listing
This opportunity was discovered on GitHub's public careers page and is republished here for discovery purposes. Applications are handled by the employer.
GitHub team? Claim this listing or ask us to remove it.