Incident Operations Specialist
Zapier $119K - $178.5K/year Posted 4 days ago
About the role
As Zapier expands into the enterprise market, incident management is a front-line function for customer trust and operational reliability. The Incident Operations Specialist is the engine that keeps the day-to-day operation of Zapier's incident program running: reliably, visibly, and at scale. Reporting to the Incident Program Manager, you own the execution of the systems, workflows, and tooling that power incident response across the company. This is not a pure engineering role; it is an ops role with technical depth, for someone who thinks in systems, writes automation to scale their own work, and leaves things more reliable than they found them.
Responsibilities
- Own incident tooling operations: maintain the reliability and configuration of incident.io, PagerDuty, Slack-based workflows, on-call rotations, and escalation paths, and fix what breaks.
- Build and maintain AI-powered workflows: incident thread summarization, postmortem draft generation, follow-up triage, severity classification, and data hygiene, turning one-off experiments into durable systems.
- Build and sustain a community of practice for Incident Commanders and Support Leads, running regular touchpoints and coaching responders on what good looks like.
- Operate data and reporting systems: build and troubleshoot dashboards and reports in Databricks, Grafana, and Looker, ensuring data quality and metric accuracy.
- Keep playbooks, templates, and incident guides current and usable under pressure.
- Participate in incidents and postmortem reviews, identify patterns, and surface recurring friction with recommended solutions.
- Coordinate across Engineering, Support, GTM, Legal, and Finance, and help onboard new stakeholder groups as the program expands.
Qualifications
- Experience in incident response, technical operations, or a reliability-adjacent role, with hands-on familiarity with incident tooling and Slack workflow automation.
- Comfortable building with SQL and reporting tools (Databricks, Looker, Grafana), and able to diagnose operational issues using logs and integration debugging.
- Able to build lightweight automations, configure API integrations, and prototype AI agents without needing an engineer to do it for you.
- Demonstrable daily AI fluency: repeatable AI-powered workflows with verification and judgment built in, and a clear account of how they changed throughput or quality.
- Strong attention to detail in configuration, data quality, and documentation, with visible, async-first working habits.
- Clear written communication across engineering, support, and leadership audiences, including translating technical incident detail for non-technical partners.
Skills
How to apply
Apply directly on the employer's application page. Your application goes straight to them.
Republished listing
This opportunity was discovered on Zapier's public careers page and is republished here for discovery purposes. Applications are handled by the employer.
Zapier team? Claim this listing or ask us to remove it.