Your new workforce won't be all human. Join the
conference to learn the new workforce OS.
request invite

Share this article

Your On-Call Now Has an AI Buddy: From Alert to a Triaged Teams War Room

Before the on-call engineer is even online, an AI buddy reads the alert, triages it, opens a Teams war room, and tags the right responder.

TL;DR: When an alert fires, an AI Coworker correlates it, opens a triaged Microsoft Teams war room, and routes it to the on-call owner before a human joins. From page to a briefed room in minutes, read-only, so responders start with context instead of four open tabs.

Related reading: How an AI SRE cut incident MTTR by 8x.

Ask anyone who carries a pager what the worst part is. It is not the fix. It is the cold start. You wake up to an alert, and before you can do anything useful you have to work out what the resource even is, whether it is actually affecting customers, who else to pull in, where the war room is, and what to check first. Fifteen minutes gone before the real work begins.

We built an AI On-Call Buddy to do that cold start for you. This is not a concept post; it is a walk through a real run on our own production environment.

The foundation: your services as data

The Buddy is only as good as what it can read, so the on-call knowledge lives in a Service custom object in Atomicwork. Each production service is a real record. For example, ATOM-VNET-PROD-01, our production Azure virtual network, is Tier 1, owned by the infrastructure team, with an on-call team, named escalation contacts, and its linked configuration items all on the record. That is the difference between an AI that guesses and one that knows who owns what.

The trigger: a real incident

An Azure Monitor Sev2 alert fires: high dropped inbound packets on ATOM-VNET-PROD-01, thousands of drops sustained over several minutes. It lands as an incident in Atomicwork with the affected asset attached. The On-Call Buddy picks it up.

What the Buddy actually did

Step by step, on the real incident: It read the ticket. The Buddy pulled the incident and its notes, then said what it saw: a Sev2 Azure packet-drop signal on a production VNet, with no confirmation yet of actual service impact or blast radius.

It triaged honestly. Rather than declare a confirmed outage, it classified the incident as investigation-needed, because impact was not yet established. And here is the part that matters most: the tooling available to it could not pull live Azure Monitor VNet metrics, so it explicitly refused to fabricate telemetry. It said so, and delivered debugging guidance instead of inventing numbers. An on-call buddy you can trust is one that tells you what it does not know.

It opened a Teams war room. The Buddy created a Microsoft Teams war room for the incident and posted a structured situation report: what and where (the alert, resource group, region, environment, fired time, and owning team); impact, observed versus potential (elevated inbound packet drops on a production VNet; potential degraded or intermittent connectivity for workloads behind it; customer-facing blast radius not yet confirmed); probable-cause hypotheses, working and unconfirmed (an NSG or firewall rule change dropping inbound flows; capacity or throughput saturation; Azure Stack HCI host or virtual switch health degradation; or an upstream Azure networking event); and debugging guidance for responders (pull the VNet dropped-packet counters and correlate against the alert time, check NSG flow logs and recent rule changes, inspect host networking and vSwitch error counters, review Azure Service Health for the region, and quantify whether dependent services are actually degraded or this is drop-counter noise).

It routed to the right person. The Buddy resolved the on-call from the roster and tagged the primary infrastructure responder by name in the war room, asking them to confirm and post first findings.

It stayed in its lane. The pinned message said it plainly: the room is for coordination and investigation only, the Buddy makes no production or configuration changes, and a human incident commander owns remediation.

Why this is the right shape for on-call AI

Two things make this work as a buddy rather than a gimmick. It does the cold start, not the fix. Everything the Buddy did, reading, triaging, war-rooming, routing, and listing what to check, is the orientation work that delays every incident. The engineer arrives to a briefed war room and a hypothesis, and spends minute one on the problem.

It is honest about its limits. When it could not see live metrics, it said so and refused to make them up. That single behavior is what makes an on-call buddy safe to trust at 2am, when you are least able to second-guess it.

The outcomes that matter

The cold start disappears. War room, situation report, responder, and debugging checklist exist before the engineer logs in. The right person, from real data. Routing comes from the Service object's owners and on-call, not a stale wiki. Trustworthy triage. Ranked, clearly-labelled hypotheses, and an explicit refusal to fabricate telemetry. Humans stay in control. Investigation and coordination from the Buddy; remediation from a named incident commander.

Frequently asked questions

What is an AI on-call buddy? An AI agent that does the orientation work at the start of an incident: it reads the alert and the affected service, triages it, opens a war room, routes to the on-call responder, and posts debugging guidance, so the human engineer starts from a briefed position instead of a cold one.

Does it resolve the incident or make changes? No. It investigates, triages, coordinates, and documents. It makes no production or configuration changes. A human incident commander owns remediation.

How does it know who is on call? It reads the Service object in Atomicwork, which holds each service's owner, on-call team, and escalation contacts, and routes accordingly. If the roster has a gap, it falls back to the owner and flags it.

What happens when it cannot get the data it needs? It says so. In the real run, it could not pull live Azure Monitor metrics, so it refused to fabricate telemetry and delivered debugging guidance for a human to execute instead.

Does it create the Teams war room automatically? Yes. It opens a Microsoft Teams war room for the incident, posts the structured situation report, and tags the primary responder.

The bottom line

On-call does not have to start cold. With the service data modeled and an AI Buddy on watch, the alert becomes a briefed Teams war room with a named responder, ranked hypotheses, and a debugging checklist, in the time it takes you to find your laptop. The Buddy does the prep and tells you what it does not know; you do the fix.

Meet 100+
tech-forward CIOs
Date icon for Atomicwork event
Sept 24, 2025
Venue icon for Atomicwork event
Palace Hotel, SF
Request an invite

Frequently asked questions

Chevron navigation icon on Atomicwork website
FAQ question text
Chevron navigation icon on Atomicwork website

You may also like...

No items found.