Skip to content

Major incidents

For the helpdesk team

When something big breaks — the MIS is down, the whole site’s lost WiFi, the email tenant is offline — a normal ticket isn’t the right tool. Major incidents give you a war-room workflow: one record, a running chronology, a stakeholder list, and a structured close.

It lives under Major incidents in the agent rail. The screen is gated by the helpdesk::major_incident::manage permission, which sits on the helpdesk agent role by default, so most agents can declare and drive an incident.

A major incident is distinct from a ticket and from a public status-page entry. It can be spawned from a ticket (the link is kept), and you can also raise a public status incident alongside it — but neither is created automatically.

Click Declare and give it a title, an optional summary, and a severity:

  • P1 — trust-wide, business-stopping
  • P2 — significant, a department or major service down
  • P3 — degraded but limping along

Declaring stamps the time, makes you the commander, and writes the first line of the chronology so the eventual review starts from a non-empty timeline.

The incident moves through a lifecycle: declared → investigating → identified → monitoring → resolved → closed.

Post updates as you go — each one lands on the timeline. An update can carry a status change in the same gesture (type “Identified — fix deploying”, flip the status, and the timeline records both as one event). Updates are typed: a status change, a comm (the default), an action taken, or the post-incident review.

Add subscribers so the right people hear from you without chasing. Each subscriber is a stakeholder, responder, or observer, and is either an internal user or an external name + email.

When you post an update, tick notify subscribers to fan it out. Each subscriber gets their own email (one per recipient, not a shared BCC) so exec phone-forwarding rules that key on the To address still fire.

Resolve marks the incident resolved and stamps the time — but it’s not terminal. You can keep posting, or step back to monitoring if it flares up again.

Close is terminal and requires a post-incident review body at the same time. There’s no closing without writing up what happened — the review is filed onto the timeline as the closing entry. That’s deliberate: the discipline of a short write-up while it’s fresh is what turns an outage into a problem record worth acting on.