Incident management 3 min

What is Incident?

An incident is an unplanned event that disrupts or degrades a service, threatens its reliability, security, or data integrity, or otherwise requires coordinated response. User impact is a primary declaration signal for many teams, but some incidents are declared before users are affected or for risks that are not immediately visible to users.

01 Mechanics

How incidents are defined and declared

Many organizations prioritize actual or imminent user impact when deciding whether to declare an incident. A failed instance that is safely absorbed by redundancy may not require an incident, while elevated user-facing errors usually do. Teams may also declare incidents for security exposure, data corruption, compliance risk, or a credible threat of impact even when users have not yet reported symptoms.

Declaration is a deliberate act with consequences. It creates a record, assigns an owner, opens a communication channel, and starts a clock. Teams that declare readily tend to respond better than teams where declaration feels heavyweight, because under-declaring means the response is informal, uncoordinated, and undocumented.

02 Value

Why the incident construct matters

The formal structure exists to make response coordinated rather than ad hoc.

  • Clear ownership: one person is accountable for driving resolution.
  • Single channel: everyone involved works from the same information.
  • Documented timeline: what happened is recorded while it happens.
  • Learning input: the record feeds the postmortem and the metrics.
03 Limits

Limits and common problems

Under-declaration is a frequent failure. Problems get handled informally in a side channel, so there is no record, no metrics, and no review, and the same failure can recur because nobody analyzed it.

Over-declaration is possible but less common, and it often points at severity levels that are poorly calibrated. Its cost is responder time and stakeholder noise, which is generally cheaper than the alternative.

04 Comparison

Incident vs outage

An outage is a specific and severe kind of incident: the service is unavailable. It is close to binary from the user's perspective, since the thing either works or it does not.

Incident is the broader category. It includes degradation that falls short of unavailability, such as elevated latency, partial feature failure, or errors affecting a subset of users, as well as security and data-integrity events that may not be user-visible at all. A system that only tracks outages will miss much of what actually needs coordinated response.

Key takeaways

  • An incident is an unplanned event requiring coordinated response, which may involve reliability, security, or data integrity.
  • User impact is a primary declaration signal for many teams, but not the only valid trigger.
  • Declaration creates ownership, a communication channel, a timeline, and a clock, which is why it should be lightweight.
  • Outages are a severe subset of incidents; much of what needs response is degradation or non-user-visible risk.

Frequently asked

Product

  • Agentic Production Engineering

Compliance

All systems normalBuilt in NYC

The autonomous system for production.
SOC 2, GDPR, and HIPAA compliant.