Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms
GitLab $126.4K - $314.4K/year Posted 4 days ago
About the role
Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale, combining software engineering with operational excellence. This is a single application for SRE opportunities across GitLab's Infrastructure Platforms department: rather than asking you to choose a team or level upfront, GitLab evaluates your skills holistically and matches you to the opportunity that fits, hiring from Intermediate through Senior Staff across multiple Infrastructure Platforms teams. You are not expected to know every technology in the environment; strong fundamentals, a growth mindset, and the ability to learn quickly matter most.
Responsibilities
- Keep user-facing services and production systems reliable, scalable, and efficient.
- Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows.
- Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling.
- Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps.
- Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately.
- Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages.
- Take part in incident response and post-incident reviews, turning learnings into changes in automation and process.
- Document runbooks, architecture decisions, and reviews so findings become repeatable practices.
Qualifications
- Experience keeping production systems reliable, combining an operations mindset with real software engineering practice.
- At Intermediate level: meaningful contributions to reliability, automation, and operational efficiency, working independently within a scoped area and diagnosing issues on your own.
- At Senior level: driving reliability improvements across multiple projects or services, leading investigations, anticipating cascading failures, and coordinating incident response.
- At Staff level: shaping reliability strategy across teams and services, defining patterns others reuse, and introducing prevention strategies for systemic weaknesses.
- At Senior Staff level: setting technical direction for reliability across a sub-department, driving the hardest systems problems, and mentoring Staff and Senior engineers.
- Strong technical fundamentals and the ability to learn GitLab's tools, systems, and ways of working quickly.
Skills
How to apply
Apply directly on the employer's application page. Your application goes straight to them.
Republished listing
This opportunity was discovered on GitLab's public careers page and is republished here for discovery purposes. Applications are handled by the employer.
GitLab team? Claim this listing or ask us to remove it.