Incident management 3 min

What is On-Call?

On-call is the rotation that guarantees a human is available to respond to production problems at any hour. One or more engineers carry responsibility for a defined window and are expected to acknowledge and act on pages during it.

01 Mechanics

How on-call is structured

A rotation assigns primary responsibility across a team, most often in weekly shifts. A secondary rotation provides backup when the primary does not acknowledge, and escalation paths reach specialists or management when the situation exceeds what the rotation can handle.

The parameters that determine whether a rotation is sustainable are its size, shift length, alert volume, and the compensation or time off that follows a disrupted night. A rotation of four people receiving frequent overnight pages will not retain those four people.

02 Value

Why on-call exists

Systems fail at all hours and users are affected regardless of business hours. On-call converts an unbounded expectation that someone will notice into a defined, compensated responsibility.

  • Guaranteed coverage: a specific person is accountable at every moment.
  • Bounded expectation: engineers not on-call are genuinely off.
  • Feedback loop: teams that operate their own software build it more carefully.
  • Response speed: someone is already prepared rather than being found.
03 Limits

Limits and sustainability problems

On-call is where operational debt is collected. Poor alerting, fragile systems, and missing automation all convert directly into pages, and the rotation absorbs the cost. A team can run indefinitely on a healthy rotation and will lose people to an unhealthy one.

Small teams face the hardest math. A three-person rotation creates a high on-call burden and may be difficult to sustain, especially with frequent pages. Below a certain team size, either the rotation spans teams or the alerting has to be quiet enough that on-call is genuinely uneventful.

04 Comparison

On-call vs an operations team

A dedicated operations team separates the people who build software from the people who run it. It provides deep operational focus, and it also creates a feedback gap: developers do not experience the consequences of what they ship.

Developer on-call closes that gap. The team that writes the service responds to its failures, which aligns incentives toward reliability. The tradeoff is that operational expertise is spread thinner. Many organizations run a hybrid, with a platform team owning shared infrastructure and product teams on-call for their own services.

Key takeaways

  • On-call converts an informal expectation into a defined, compensated responsibility with guaranteed coverage.
  • Sustainability depends on rotation size, shift length, alert volume, and recovery time after disrupted nights.
  • It is where operational debt is collected, since poor alerting and fragile systems convert directly into pages.
  • Developer on-call aligns build and run incentives, at the cost of spreading deep operational expertise thinner.

Frequently asked

Product

  • Agentic Production Engineering

Compliance

All systems normalBuilt in NYC

The autonomous system for production.
SOC 2, GDPR, and HIPAA compliant.