Service Incident
Pronunciation: SUR-vis IHN-suh-dunt
Definition
A service incident is an unplanned interruption, degradation, error, or control failure that affects delivery or quality of a service. Service Incident should distinguish an alert, suspected event, confirmed incident, material impact, and restored service because each state requires different decisions and notifications. Service Incident must define the affected service or asset, event severity, business and customer impact, evidence, responsible roles, containment priority, recovery objective, and reporting obligations.
Overview
Service incidents include outages, latency, incorrect results, unavailable dependencies, failed settlements, data inconsistency, capacity exhaustion, and customer-facing errors. Security events may cause service incidents, but operational or vendor failures can produce similar impact.
Restoring availability does not always restore correct business state. Payments may remain duplicated, delayed, unreconciled, or incorrectly communicated after technical recovery, while retries can worsen damage when idempotency or settlement status is unclear.
Teams should classify impact, assign command, communicate verified facts, stabilize service, reconcile affected records, and meet contractual or regulatory duties. Closure requires confirmed recovery, customer remediation, evidence, root-cause analysis, and tracked prevention or resilience improvements. Incident metrics should include affected transactions and unreconciled value, not availability alone.
Unlike a security incident, a Service Incident is used for any material service degradation or interruption; for example, a dependency outage can require coordinated recovery even when no compromise occurred.
A service incident is an unplanned interruption, degradation, error, or control failure that affects delivery or quality of a service. Service Incident should distinguish an alert, suspected event, confirmed incident, material impact, and restored service because each state requires different decisions and notifications. Service Incident must define the affected service or asset, event severity, business and customer impact, evidence, responsible roles, containment priority, recovery objective, and reporting obligations. Service-incident response must restore both technical function and correct business state, including reconciliation, communication, remediation, and prevention.
A production treatment of Service Incident should test an unplanned interruption, degradation, error, or control failure that affects delivery or quality of a service within the relevant asset, decision, or service state. The Service Incident context record for unplanned interruption, degradation, and error should preserve source data, configuration or policy version, responsible actor, exception, and outcome. Review of Service Incident should determine whether safeguards addressing unplanned interruption, degradation, and error changed exposure in practice, not merely whether a document or setting existed.
Key Takeaway
Service-incident response must restore both technical function and correct business state, including reconciliation, communication, remediation, and prevention.
Sources
- NIST Documentation: Cyberframework — NIST (2026-07-30)