How to Automate PagerDuty

PagerDuty is already one of the most automated products a team owns. Event Orchestration routes and enriches alerts, escalation policies find a human, and Incident Workflows run the response. The gap is outside it: the vendor status page, the admin console and the ticket queue that say whether this incident matters.

PagerDuty decides who wakes up

PagerDuty is the on-call and incident response platform sitting between everything that monitors a system and the people who have to do something about it. Alerts arrive from dozens of sources, get turned into incidents, and get put in front of whoever is on the rota at that hour.

Its real job is smaller and more serious than the feature list suggests: deciding who gets woken up, and making sure that if the first person does not answer, the second one does. Everything else in the product exists to make that decision correct more often.

It has also spread well beyond engineering. Hospitals route unacknowledged telemetry alarms through it. Biomedical engineering teams escalate equipment checks that went past due. Facilities and workplace teams use it for anything where an unanswered message is a real problem rather than an annoyance.

Wherever it is installed, it holds the schedule, the escalation path and the record of who responded.

The first ten minutes go on gathering context

The page fires at two in the morning. Somebody is awake and acknowledging inside ninety seconds, which is the part that works.

Then the actual work starts, and almost none of it happens in PagerDuty. What deployed in the last hour. Whether the cloud provider or the payment processor has anything on their status page. Whether this same alert fired last Tuesday and turned out to be nothing. Which customers are on the affected instance, and whether one of them is the account with a renewal call at nine. Whether support has already had three tickets about it, which would mean it is real.

Every one of those is a different login, at two in the morning, by somebody who has been awake for four minutes.

The same shape repeats after the incident, in daylight. The writeup that needs the timeline reassembled. The ticket that never got updated. The action items everybody agreed to and nobody scheduled.

The pagerduty.com homepage, the app these three jobs run in. PagerDuty
Context gathered at page time The vendor status pages, the recent changes and the open tickets checked in the first minute, while the responder is still finding their laptop.
Blast radius named Which accounts sit on the affected instance, read from the systems that know, so stakeholder updates go to a list rather than to everyone.
Follow-ups that do not quietly expire The action items agreed after the incident, checked a fortnight later, with the ones nobody closed put back in front of someone.

Event Orchestration acts on the events it is sent

It is worth being straight about this: PagerDuty is heavily automated already, and more so than most of the tools it connects to. Anyone expecting a manual product will be disappointed.

Event Orchestration applies nested rules to incoming events, so alerts get routed, enriched, suppressed or escalated before an incident exists at all. Escalation policies and on-call schedules move a page along until a human answers, with overrides and handoffs handled. Incident Workflows run a defined response when something is declared: open the channel, page the right team, update the status page. Automation Actions and Runbook Automation let a responder run a diagnostic or an approved runbook from inside the incident. On top of that, AIOps noise reduction groups and correlates related alerts so one problem does not arrive as four hundred pages.

The ceiling is not capability, it is reach. All of it acts on events that were sent in, or runs jobs somebody already built and connected. Event Orchestration can enrich an alert with what is inside the alert.

The things that would tell you whether this matters are on other people's websites and in systems that have never sent PagerDuty an event.

The page can arrive with the context attached

The first ten minutes of an incident go on looking up things that were already published somewhere.

The status pages of the four vendors you depend on, checked and quoted in the incident channel within a minute of the page, so nobody spends twenty minutes debugging someone else's outage. The list of accounts on the affected instance, read from the admin console, so the stakeholder update names who is actually affected. The support queue checked for customers already reporting it. The change log for the last two hours, sitting there before anyone asks.

Outside engineering it works the same way. An unacknowledged alarm escalated with the patient's unit and who is genuinely on shift, rather than who the rota says. An overdue equipment check escalated with that asset's service history from the maintenance system attached.

Then the tail end. Action items from the review checked a fortnight later, with anything still open listed by owner. The ticket that should have been updated, flagged while somebody still remembers the detail.

Reading and assembling can run unattended. Waking a second team, posting to a public status page or telling customers anything should still be a person's call.

The responder should be reading, not fetching

In the middle of the night, the vendor's status page is one more tab to open on top of working out what actually broke. Opening it and diagnosing the outage are different jobs that happen to have landed on the same person.

WebRun is an agent that works a real Chrome browser, signed in as you. When an incident opens it goes and looks: the status pages, the admin console, the ticket queue, the deploy history, and whatever else your runbook says to check first. It posts what it found where the response is happening.

It runs in your own private environment and you can watch a run and stop it. It does not resolve incidents, page teams or publish anything customers see. Those stay with the person holding the incident.

The workflows below are already built, and each one names what it opens and what it reports.

Questions people ask

PagerDuty already automates a lot. What is actually missing?

Reach, not capability. Event Orchestration and Incident Workflows act on events that were sent in and on tools that were wired up. The vendor status page, the customer admin console and the ticket queue send nothing, and those hold the context that decides how serious the page is.

Will it resolve or acknowledge incidents by itself?

No. Acknowledging is how a team knows a human has it, so an agent doing that would break the one thing PagerDuty exists to guarantee. It gathers and reports. People acknowledge, escalate and resolve.

Can it work with our existing escalation policies and schedules?

Yes, because it does not replace any of it. Your schedules, policies and orchestration rules keep running exactly as configured. This adds the looking-around step that currently happens in a responder's browser at two in the morning.

13 ready-made PagerDuty workflows

Each one names the apps it touches and the exact steps it takes. Open one to read what it will do, then turn it on.

Automated Unacknowledged Alarm Escalation
WebRun watches Norav Medical for alarms sitting unacknowledged past your threshold, pages the next person in the on-call chain through PagerDuty, and never lets one go quiet.
Norav MedicalPagerDutySlack
Automated NRC Health Rounding Action Item Reminders
WebRun texts the owner of any overdue rounding action item through Twilio each morning, and escalates it through PagerDuty to the unit director's on-call if it's still open after 72 hours.
NRC Health RoundingTwilioPagerDuty
Automated Nuvolo Overdue PM Escalation Alerts
Every morning, WebRun opens Nuvolo, finds preventive maintenance tasks that are now past their grace period, pages the on-call supervisor through PagerDuty, and emails a daily escalation summary to your biomedical engineering manager.
NuvoloPagerDutyOutlook
Automated NRC Health Service Issue Follow-Up
WebRun checks NRC Health Rounding hourly for service issues logged during rounds, posts each open item to the responsible department's Microsoft Teams channel, and escalates through PagerDuty if it's still open past your SLA.
NRC Health RoundingMicrosoft TeamsPagerDuty
Automated NRC Health Critical Feedback Escalation
When a round captures a critically low score or comment, WebRun pages the on-call patient experience leader through PagerDuty and logs the case in an internal Outlook email so nothing falls through.
NRC Health RoundingPagerDutyOutlook
Automated Envoy Emergency Roll Call Drafts
The instant a fire warden starts it, WebRun pulls everyone currently checked into Envoy, pages the on call safety officer through PagerDuty, and drafts a roll call list by floor in Slack for the warden to call out at the assembly point.
EnvoyPagerDutySlack
Automated Escalation Response Time Report
WebRun pulls escalation and acknowledgment timestamps from Norav Medical and PagerDuty each week, calculates response times, and posts the report to Slack and Google Sheets.
Norav MedicalPagerDutyGoogle Sheets
Automated NRC Health Call Light Delay Alerts
WebRun logs every call light delay a patient mentions during rounding into Excel and pages the nursing supervisor on-call through PagerDuty when a unit crosses your threshold in a single shift.
NRC Health RoundingExcelPagerDuty
Check Critical User Journeys Hourly
Every hour, WebRun runs through your most important user flows - sign up, checkout, login - and alerts the team the moment a step breaks.
Your appDatadogPagerDuty
Summarise Incident Timelines Automatically
After an incident is resolved, WebRun reads the alert history, Slack thread, and ticket trail and writes a structured timeline so your post-mortem starts with the facts already assembled.
PagerDutyJiraSlack
Monitor Vendor Status Pages for Outages
WebRun checks your critical third-party vendor status pages every few minutes and fires a Slack alert the moment a vendor reports degraded performance or an outage.
StatuspagePagerDutySlack
Write Your On-Call Handoff Summary
At the end of every shift, WebRun reads the week's incidents and alerts, writes a structured handoff report, and sends it to the incoming on-call engineer before they pick up the pager.
PagerDutyJiraSlack
Run Synthetic Uptime Checks
Every few minutes, WebRun hits your critical endpoints, measures response time, and pages the on-call engineer the moment anything goes down.
Your appPagerDutySlack

Want one of these running on your own PagerDuty?

Show WebRun the process once and it will run it on schedule, in your own private browser environment.