On-Call
Alerts and incidents
An alert is one problem reported by a monitoring tool (or raised by hand). An incident is a coordinated response your team declares, with a timeline, a public status page entry and a postmortem. This page covers both.
Alert statuses
| Status | Meaning |
|---|---|
| Triggered | New and unanswered. The escalation policy is running. |
| Acknowledged | Someone is on it. The escalation stops (or pauses, if the policy has an acknowledgement timeout). |
| Resolved | Over. Resolved by a person, or automatically by the source's recovery notification. |
| Suppressed | Silenced from a voice call (key 6). Escalation stops. |
"Open" means triggered or acknowledged.
The Alerts page
On-Call → Alerts lists your organization's alerts.
- Counters at the top show Total, Triggered, Acknowledged and Resolved.
- The status filter offers All, Open (the default), Triggered, Acknowledged, Resolved and Suppressed.
- Assigned to me narrows the list to alerts assigned to you; it combines with the status filter.
- Search title or source… searches as you type.
- Filters adds Severity, Source, Assignee, Escalation policy, Integration and Created between.
- Expand a row to see its description, custom details (labels) and fingerprint, or click View full alert.
Each integration's detail dialog also has View this integration's alerts, which opens this list filtered to that integration.
Respond to an alert
Open an alert to see its details, labels and timeline. The actions available depend on its status and your permissions:
| Action | What it does |
|---|---|
| Acknowledge | Marks the triggered alert as acknowledged and stops (or pauses) the escalation. |
| Take over | Assigns the alert to you, acknowledges it if needed and stops the escalation. See Overrides and take over. |
| Redirect | Hands the alert to another escalation policy. Its current escalation stops and the selected policy is paged from the top. An acknowledged alert goes back to triggered so the new policy actually pages. |
| Resolve | Closes the alert and stops the escalation. |
| Add a comment… | Adds a note to the alert's timeline under your name. |
Acknowledging an alert that is already acknowledged, or resolving one that is already resolved, does nothing and is not an error. A resolved alert cannot be acknowledged again.
Responding (acknowledge, take over, resolve, notes) needs permission to respond to alerts. Redirecting and creating alerts by hand need permission to write alerts. Owners, Admins and Members have both by default.
Bulk actions
Tick alerts on the Alerts page (or Select all open alerts) to act on many at once: Acknowledge N (triggered alerts), Take over N (open alerts) and Resolve N. Each alert is processed on its own, so one that cannot be changed does not stop the rest.
The alert timeline
Every alert keeps a timeline of what happened to it, including:
- Alert created, Acknowledged, Resolved, Taken over by …, Redirected to another policy
- Notifications sent: Voice call to …, Push notification to …, Email sent to …, and the call result (for example Not answered or Line busy)
- Escalated to next step when someone pressed 3 on a call
- No one was on call — this step reached nobody
- Notifications held back by a person's notification schedule, by a maintenance window, or because your organization has used its free allowance with no payment method on file
- Retriggered ×N when the source sent the same alert again
- Comments, including those added from Slack
When someone asks "why wasn't I paged?", the alert's timeline is the place to look.
Deduplication and auto-resolve
Every alert has a fingerprint that identifies the problem. Integrations derive it from the source's own identifiers (for example, Alertmanager's fingerprint, a Zabbix host and trigger, or a PRTG sensor ID).
- If an alert arrives while an alert with the same fingerprint is still open (triggered or acknowledged), EvoHub does not open a second one. It records Retriggered on the existing alert instead, and nobody is paged again.
- Once the alert is resolved, the next alert with that fingerprint opens a new alert.
- When the source reports recovery, EvoHub resolves the open alert with the matching fingerprint. Each integration page says how its source reports recovery.
Create an alert by hand
Click New Alert on the Alerts page, enter a Title, choose a Severity (Critical, High, Medium, Low or Info), optionally choose an escalation policy and add a Description. If you choose a policy, it starts paging immediately — this is a convenient way to test a policy end to end.
Maintenance windows
While a maintenance window is in progress, On-Call pages nobody in your organization. Alerts are still recorded, and each skipped notification shows on the alert's timeline as suppressed by maintenance. Schedule windows in On-Call → Maintenance (Schedule maintenance); a window can also be shown on a status page and mute uptime monitors. See Incidents and maintenance on status pages and Uptime alerts and silence windows.
Incidents
An incident is a record your team declares for a significant problem. Incidents do not page anyone by themselves; they are where you coordinate, communicate and learn.
Declare an incident
Go to On-Call → Incidents → New Incident and enter:
- Title and an optional Summary (you can edit both later).
- Severity: Critical, Major, Minor or None.
- It already happened: tick this to record an incident that has already ended, with its own Started and Ended times. Nobody is paged, and if you publish it to a status page it shows on the days it happened.
Incidents are numbered (INC-1, INC-2, …).
Work an incident
An incident moves from investigating to identified (Mark Identified) to resolved (Resolve). On the incident page you can:
- Edit the title, summary and severity.
- Add timeline notes (Add a timeline note…), and edit or delete them later.
- See the Commander and when it Started.
Publish to a status page
Click Publish to status page, choose the Status page, the Public impact (None, Minor, Major, Critical), the Affected components, and an optional Public message. Once published, the incident shows On status page, and EvoHub keeps the two in sync: marking it identified or resolved, and adding, editing or deleting timeline notes, update the public incident too. Unpublish removes it from the status page.
Publishing an incident that has already ended puts it on the page's history on the days it happened; subscribers are not emailed. You need a status page first — see Status pages overview.
Postmortem
After an incident is resolved, click Postmortem to write one. Start from a template (Standard, 5 Whys or Brief) and fill in Summary, Root cause, Contributing factors, Timeline, Action items and Lessons learned. Save internal keeps it inside EvoHub; Publish publicly shows it on the incident's public status page entry (when the incident is published there).
Related
Was this page helpful?
