Zapier Logo

Zapier

Incident Operations Specialist

Reposted 20 Days Ago
Remote
Hiring Remotely in United States
119K-179K Annually
Mid level
Remote
Hiring Remotely in United States
119K-179K Annually
Mid level
Run and maintain Zapier's incident program operations: manage incident tooling and on-call systems, build AI-powered automations and dashboards, maintain runbooks and enablement, operate reporting and observability, coordinate cross-functional stakeholders, and drive continuous improvement to keep incident response reliable and scalable.
The summary above was generated by AI
AI at Zapier

At Zapier, we build and use automation every day to make work more efficient, creative, and human. So if you’re using AI tools while applying here - that’s great! We just ask that you use them responsibly and transparently.

Check out our guidance on How to Collaborate with AI During Zapier’s Hiring Process, including how to use AI tools like ChatGPT, Claude, Gemini, or others during our hiring process - and when not to.

 

Job Posted: July 10th, 2026

Location: NAMER

Hi there!

As Zapier expands into the enterprise market and accelerates AI-driven development, incident management is increasingly critical to customer trust and operational reliability. The Incident Operations Specialist keeps that program running day to day through reliable tooling, clean data, repeatable workflows, and AI-powered automation.

You’ll report to the Incident Program Manager and help shape how Zapier responds to incidents, learns from them, and supports the people doing that work. This is an operations role with technical depth, not a software engineering role. We care most about two things: proven incident management experience and genuine AI fluency. The rest is coachable.

  • Our Commitment to Applicants

  • Culture and Values at Zapier

  • Zapier Guide to Remote Work

  • Zapier Code of Conduct

  • Diversity and Inclusivity at Zapier

What we're looking for

AI fluency (required, not optional). This is a hard requirement at point of hire, not something you'll grow into on the job. Concretely, we're looking for:

  • You use AI-native tools (Cursor, Claude, Copilot, or similar) as your default working environment, not as a novelty.

  • You've built AI-powered workflows that keep running when you're offline, not one-off prompts. You can describe two or three specific examples, what they replaced, and what verification you built in.

  • You can quantify how AI has changed your throughput or quality.

  • You know when AI output needs checking, especially under incident-time pressure, and you have a point of view on how you calibrate trust.

  • Please note: If your AI usage is mostly occasional prompting of a chat interface, this role isn't the right fit yet.

Incident response and analysis experience: You've worked in incident response, reliability, or a closely adjacent role. You've been hands-on with tools like incident.io and PagerDuty: on-call rotations, escalation paths, routing, integrations. You've analysed incidents after the fact, spotted patterns across many of them, and turned that into program-level improvements.

Technical depth to build your own tools: You're not a software engineer, but you can write SQL against Databricks, wire up API integrations, build Slack workflows, and prototype lightweight AI agents. If a workflow doesn't exist, you build it. If a dashboard is broken, you fix it.

How you work: Async-first and visible: status in public channels, no need to chase. You close the loop, prioritise ruthlessly, and push back on off-program requests rather than getting pulled thin. You translate technical detail into plain language for Support, GTM, and leadership without losing the signal.

What you'll do
  • Own incident tooling operations. Keep incident.io, PagerDuty, on-call rotations, escalation paths, and Slack-based workflows configured, reliable, and integrated. Fix what breaks.

  • Build and maintain AI-powered workflows. Thread summarisation, postmortem drafting, follow-up triage, severity classification, data hygiene. Turn one-off experiments into durable systems.

  • Analyze incidents and drive improvement. Participate in incidents and postmortems, spot recurring patterns, surface program-level friction with recommended fixes, not just problems.

  • Operate data and reporting. Build and troubleshoot dashboards and reports (Databricks, Grafana, Looker). Guard data quality and metric accuracy.

  • Sustain the IC community. Grow the community of practice for Incident Commanders and Support Leads. Coach responders on what good looks like.

  • Keep documentation usable under pressure. Playbooks, templates, guides. Flag gaps where program-level guidance needs updating.

Our stack
  • Incident: incident.io, PagerDuty, Slack

  • Data and observability: Databricks, Grafana, Looker, SQL, Datadog, Prometheus, Opensearch, Graylog

  • AI: Cursor, Zapier AI, Claude, or equivalent

  • Collaboration: GitLab, Coda, Google Workspace, Jira, Zendesk

Application Deadline:

The anticipated application window is 30 days from the date job is posted, unless the number of applicants requires it to close sooner or later, or if the position is filled.

Even though we’re an all-remote company, we still need to be thoughtful about where we have Zapiens working. Check out this resource for a list of countries where we currently cannot have Zapiens permanently working.

Similar Jobs at Zapier

4 Days Ago
Remote
United States
192K-287K Annually
Senior level
192K-287K Annually
Senior level
Artificial Intelligence • Productivity • Software • Automation
Lead IT as a product for a ~800-person remote company: build AI-first service delivery, architect enterprise-grade identity and device trust (Okta, Jamf), govern a 200+ app SaaS portfolio, own major incident response, partner with Security on zero-trust and compliance, and scale a globally distributed IT team while enabling enterprise sales.
Top Skills: AbacAWSEdrInfrastructure As CodeIso 27001JAMFMfaOidcOktaRbacSAMLScimSoc 2 Type 2SsoZapier
5 Days Ago
Remote
United States
192K-287K Annually
Senior level
192K-287K Annually
Senior level
Artificial Intelligence • Productivity • Software • Automation
The role entails forecasting, planning, and analysis for Zapier's self-serve and new products businesses, requiring strong SQL skills and finance expertise.
Top Skills: SQL
9 Days Ago
Remote
United States
308K-463K Annually
Senior level
308K-463K Annually
Senior level
Artificial Intelligence • Productivity • Software • Automation
Lead Zapiers security strategy for an AI-native SaaS platform. Own App and Infrastructure Security, Detection & Response, GRC, risk management, incident response, bug bounty triage, and enterprise security enablement. Partner with Product, Engineering, Legal, and GTM to embed secure-by-default design, shape AI-specific controls, support enterprise deals, and build a high-performing security engineering organization.
Top Skills: AIBug Bounty PlatformsCloudDetection & ResponseForensicsIdentity And Access Management (Iam)Incident ResponseSecure SdlcThreat Modeling

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account