Outage & service restoration
Keep customers informed through an outage without melting the queue
An incident inverts your contact centre: volume multiplies in minutes, every caller wants the same three facts, and the callers who genuinely need priority handling are buried in the queue. This blueprint absorbs the volume with location-aware updates drawn from the system of record, and makes hazard and vulnerability signals jump the queue rather than wait in it.
- The team
- 3 AI worker roles escalating into 1 human decision owner.
- The process
- 6 orchestrated stages, each with a named decision owner.
- The channels
- Voice, SMS, Web chat, WhatsApp on one shared context.
The shift
What breaks today, and what changes
The challenge
- Incident volume arrives in minutes, so the queue is already unmanageable before staff can be redeployed.
- Customers get inconsistent restoration estimates depending on who they reach and when.
- Hazard reports and medically vulnerable customers wait in the same queue as routine status questions.
What changes
- Every caller and chat gets the current, location-specific position from the system of record.
- Hazard language and vulnerable-customer flags bypass self-service and reach the right team immediately.
- Restoration updates and confirmations go out proactively, which removes the second and third contact.
The team
Every seat has one job and a defined level of autonomy
This is what gets configured in Team Canvas: role-scoped AI workers, the humans they escalate into, and a trust level on each connection that sets what a worker may do on its own. Trust starts tight and widens on evidence.
AI workers
Incident update worker
Confirms location, matches it to a known incident, and shares the approved current position.
- Owns
- Customer communication
- Trust level
- Acts within policy — may only communicate estimates from the incident record
Safety triage worker
Listens for hazard language, medical dependency, and critical-service need on every turn.
- Owns
- Priority classification
- Trust level
- Acts within policy — escalation is mandatory on any hazard signal
Field liaison worker
Turns actionable reports into structured jobs for operations with location and access detail.
- Owns
- Operational handoff
- Trust level
- Acts with approval — creates jobs, never dispatches crews
Human decision owners
Control room operator
Owns dispatch, prioritisation, and every safety decision during the incident.
Owns: Dispatch and safety authority
The process
The conversation starts the work. The workflow finishes it.
Each stage records the decision that was made and the role accountable for it — which is what makes the run reviewable afterwards rather than a black box. Where the platform executes a stage rather than a person or agent performing it, that is named too, and accountability still sits with the role.
- 1
Locate
Is this address inside a known incident, or is it an isolated fault?
Incident update worker
- 2
Assess risk
Any hazard, medical dependency, or critical-service exposure? If yes, leave the flow now.
Safety triage worker
- 3
Inform
What is the approved current position for this location right now?
Incident update worker
- 4
Create job
Does this need a field visit, and does operations have what it needs to act?
Field liaison worker
- 5
Notify on change
Has the estimate or state changed enough to proactively tell affected customers?
Incident update worker
Run by workflow
- 6
Confirm restoration
Is service actually restored for this customer, or is a follow-up needed?
Field liaison worker
Run by workflow
Decision rules
The non-negotiables encoded in the workflow, not left to a prompt.
- Gas, electrical, fire, and injury language bypasses self-service immediately and unconditionally.
- Restoration estimates are only ever repeated from the incident record — no worker computes or softens one.
- Medically dependent and critical-service accounts are prioritised automatically from the account flag.
- Every incident communication is logged against the incident for post-event review.
Controls & guardrails
What makes this safe to run in production, and provable afterwards.
- Unconditional emergency routing that no conversation path can override.
- Source-controlled estimates traceable to the incident record.
- Automatic priority flags for vulnerable and critical-service accounts.
- Complete incident communication log for regulatory review.
Platform capabilities
What this blueprint runs on
Nothing here is bespoke. Each blueprint is a configuration of the same platform, which is why the second one you launch reuses the governance you already reviewed.
Channels
- Voice
- SMS
- Web chat
Systems it reads and writes
- Outage management system
- CIS / billing
- Field service management
- Mass notification gateway
Measurement
What we instrument from day one
These are the measurements the pilot puts in place, not benchmark results. You set the targets against your own baseline — and the same instrumentation is what decides whether a worker's trust level widens.
- Surge absorption
- Measures concurrent contacts resolved during an incident window against the pre-incident baseline.
- Priority latency
- Times how long a hazard or vulnerability signal takes to reach the responsible team — the metric that matters most during an incident.
- Estimate consistency
- Confirms every communicated estimate traces to the incident record, so no two customers get different answers.
Contacts handled at peak
Hazard to dispatch
One source of truth
Rollout
How this one goes live
A deliberately narrow start, supervised, with autonomy widened per role once the record supports it.
Weeks 1–3
Status only, hazard routing live
Handle location-based status questions during normal operations. Hazard triage is enabled from day one — it is the control, not a later feature.
Weeks 4–6
Add proactive notification
Switch on outbound updates when the incident record changes, which is what removes repeat contacts during an event.
Incident-ready
Full surge configuration
Run a rehearsed incident scenario, confirm priority latency against target, then enable the blueprint as your standing surge capacity.
Questions we get asked
Outage & service restoration, in practice
- What stops a customer with a gas leak from being stuck in self-service?
- A dedicated safety triage worker evaluates every turn for hazard language, and the escalation is mandatory rather than discretionary — no conversation path, intent match, or containment target can override it. Priority latency from hazard signal to responsible team is instrumented and reviewed.
- Where do restoration estimates come from?
- Only from your outage management system. The workers repeat the approved current position for the caller’s location and are not permitted to compute, interpolate, or soften an estimate, which is what keeps two callers from getting different answers.
- How does it know which customers are vulnerable?
- From the flags already on the account in your CIS. Those flags drive automatic prioritisation in the workflow, so a medically dependent customer does not have to explain their situation to reach the front of the queue.
What changes in Energy & utilities
Incident spikes are handled with location-aware updates from the system of record, and hazard or vulnerability signals jump the queue.
Read the full Energy & utilities analysisRelated blueprints
Teams that run next to this one
Next step
Build this team around one measurable win.
We map your version of this process, compose the worker roles, connect your systems, and agree the escalation points with the people who own them — then run it supervised.