Incident Management
How Incident Management works: alert sources, routing, teams, on-call schedules, escalation, notifications, status pages and reports, and how to set it up.
Incident Management (IM) turns alerts into incidents and makes sure the right person is woken up. It works next to Monitoring: Uptimeify's own checks are one alert source among others (Zabbix, Datadog, Grafana, Prometheus, Sentry, email, your own webhook or code).
Everything below can also be done through the Incident Management API.
Who can use it
Organization members with the role admin, editor or responder. Customer logins (readonly) have no access. Within a team, every member (team role admin, member or stakeholder) may work on that team's incidents; team admins may also change the team itself. Organization admins may do both for every team.
Setting it up
- Activate Incident Management for your organization (organization admin). In the dashboard: Incidents. API: Activate.
- Create a team and add its members (Incidents → Teams).
- Create a schedule for the team: who is on call when, and how the rotation moves on.
- Build the escalation chain of the team: tier 1 is paged first; each tier has its schedules and decides when the incident moves on to the next tier.
- Connect an alert source (Incidents → Sources): pick a preset, copy its ingest URL into your tool. The setup guides cover each tool.
- Set your own notifications (Incidents → Notifications): how and after how long you are reached, and verify your phone number for SMS and voice calls.
- Send a test page to check that paging reaches you.
From alert to incident
- An alert is one signal from a source. Alerts with the same deduplication key belong together, so a flapping check does not open ten incidents.
- An incident groups alerts and is what people work on. It has a severity (
sev1most severe tosev4) and a status:triggered(nobody has reacted yet),acknowledged,investigating,identified,monitoring, and finallyresolved(ormergedinto another incident). - Routing rules decide which team a new incident goes to, for example by source or customer, and can override its severity. Without a matching rule, the source's default team takes it.
- The monitoring bridge (organization settings) brings incidents from Uptimeify's own checks into Incident Management automatically.
On-call and escalation
- A schedule defines rotations: who is on shift and when the shift hands over.
- An override replaces the schedule for a period: someone takes over a shift, or hands it off while unavailable.
- The escalation chain pages tier 1 first. If nobody acknowledges within the tier's time, the next tier is paged. A tier can be paged again before moving on, and the whole chain can repeat.
- On an incident you can acknowledge (stops escalation), snooze (pause escalation and reminders for up to 7 days), escalate now (page the next tier immediately), add a responder, comment, change severity, merge duplicates and resolve.
How you are notified
Each person decides for themselves how they are reached:
- Channels: push (mobile app), SMS, voice call, email. SMS and voice need a verified phone number.
- Rules per urgency: for example "push immediately, SMS after 5 minutes, call after 10 minutes".
- Severity filter: for example no SMS or calls for
sev4. - The notification log shows every notification that was sent, skipped or failed for you.
Shared channels (webhook, Slack, email) belong to the organization or a team. One of them can be the fallback channel: the last resort when an escalation runs out of tiers.
Details: Personal Notification Settings, Channels.
Status pages and other tools
- Status page rules switch a status page to warning or degraded while a team has open incidents of a given severity, with a manual override.
- Outbound integrations forward incident events to Slack, Microsoft Teams, Discord, Jira, PagerDuty, Opsgenie or a generic webhook. Failed deliveries can be inspected and retried.
Reports and audit
- Analytics: incident volume, MTTA and MTTR per day, the noise ratio of alerts to incidents, and on-call load per person.
- Reports: incident report (by team, severity or month) and on-call report (minutes per person), also as CSV.
- Audit log (organization admins): who changed what in Incident Management.
API
Every step above has an endpoint: Incident Management API. Most endpoints accept an organization-wide API token; personal on-call status and notification settings need a user session.
Incidents
Incidents are created automatically when monitoring detects a problem with a website or service. They are the foundation for alerting, reporting, and the public Status Pages.
Maintenance
Maintenance windows let you plan work (deployments, updates, migrations) without triggering noisy alerts. During an active window the affected target is treated as in maintenance. Alerts are suppressed and, if the customer has a status page, the service is shown as Maintenance instead of Degraded.