Staff Software Engineer Cloud Foundations
GitHub $140.4K - $372.3K/year Posted 4 days ago
About the role
The Cloud Foundations team owns and operates the foundational compute infrastructure that powers GitHub, today and through its move to Azure. The team is responsible for the systems other engineering teams build on every day: the internal VM platform, fleet lifecycle automation, configuration management, base OS and container image pipelines, and the fleet inventory and discovery layer that ties it all together. As a Staff Software Engineer, you will lead the technical direction of a platform that hundreds of internal teams depend on, building the cloud-native platform GitHub runs on next while safely evolving and retiring the data center services that came before.
Responsibilities
- Lead technical decision making and architecture across the Cloud Foundations surface area, including compute lifecycle, configuration management, infrastructure orchestration, and fleet inventory.
- Design and implement scalable, reliable solutions for complex problems such as fleet-wide reboot orchestration, immutable image pipelines, multi-region state management, and service discovery at scale.
- Define and build GitHub's Azure paved paths, including event-driven compute and cloud-native patterns.
- Drive the evolution of foundational systems from data center to Azure-native equivalents.
- Write, review, and maintain code primarily in Go, partnering with the team on language and tooling choices for new systems.
- Mentor other engineers in technical and architectural decision making, raising the engineering bar through code review and design feedback.
- Participate in on-call rotations, leading incident response on hard problems and driving follow-up work that prevents repeat issues.
Qualifications
- 9+ years of software engineering experience delivering production software in languages such as C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python (less with a relevant degree, or equivalent experience).
- 4+ years building and supporting large, high-traffic applications at scale in platform or infrastructure domains.
- 4+ years supporting and building cloud-native workloads in Azure, AWS, or Google Cloud.
- 2+ years operating fleet-scale compute infrastructure such as hypervisor platforms, configuration management systems, or orchestration tooling.
- Preferred: experience designing or operating paved-path platforms on Azure (Functions, App Service, AKS) or comparable offerings, and configuration management (Puppet, Chef, Ansible) at fleet scale.
- Preferred: experience building or maintaining planetary-scale engineering systems within a remote, distributed team.
Skills
How to apply
Apply directly on the employer's application page. Your application goes straight to them.
Republished listing
This opportunity was discovered on GitHub's public careers page and is republished here for discovery purposes. Applications are handled by the employer.
GitHub team? Claim this listing or ask us to remove it.