Payment Provider Outage
Pronunciation: PAY-munt pruh-VY-dur OW-tij
Also known as: PSP Outage
Definition
Payment Provider Outage is a period when an external payment provider cannot deliver one or more contracted payment functions for the affected merchants, regions, methods, or operations. The provider may reject requests, time out, stop sending events, return stale status, or fail to reach an underlying rail even when the merchant platform is healthy. Its boundary matters because it is attributed to a provider dependency, whereas a rail outage affects shared payment infrastructure and a service outage describes the user-facing service result. Payment teams should detect with independent probes, classify affected functions, and stop unsafe retries and retain evidence that supports recovery, investigation, and reconciliation.
Overview
Payment Provider Outage is a period when an external payment provider cannot deliver one or more contracted payment functions for the affected merchants, regions, methods, or operations. The provider may reject requests, time out, stop sending events, return stale status, or fail to reach an underlying rail even when the merchant platform is healthy. Useful measures include provider availability, affected count and value, detection time, failover success, unknown-state duration, recovery time, and post-incident exceptions.
Its boundary with Payment Provider Risk must remain explicit so related records do not collapse into one status. The record should retain provider, service and region, start and end times, affected requests, timeout and error evidence, provider notices, route changes, retries, and final outcomes.
Payment Provider Outage should remain distinct from Payment Provider Risk, Payment Rail Outage, and Payment Route Candidate, because each can represent a different stage, record, control, or financial outcome.
Governance for Payment Provider Outage needs a named owner, review cadence, approved changes, and rollback. Important failure modes include noisy alerts, blind spots, stale dashboards, missing ownership, incorrect uptime calculations, slow escalation, and recovery claims that are not verified against payment outcomes.
Production use requires a named owner, affected population, and lifecycle boundary. Controls should detect with independent probes, classify affected functions, stop unsafe retries, query authoritative status, activate approved alternatives, preserve evidence, and reconcile recovery. When Payment Route Candidate is involved, the link must be auditable so operators can decide whether retry, repair, return, rerouting, or adjustment is safe. Monitoring should define scope, measurement window, threshold, severity, owner, evidence, escalation path, and the recovery condition that closes the alert or incident.
Key Takeaway
For Payment Provider Outage, teams should detect with independent probes, classify affected functions, and stop unsafe retries, preserve authoritative evidence, and monitor provider availability, and affected count before treating the related payment outcome as complete.
Sources
- Google SRE: Monitoring Distributed Systems — Google (2026-08-03)
- NIST SP 800-61 Rev. 3: Incident Response — National Institute of Standards and Technology (2026-08-03)
- OWASP Third-Party Payment Gateway Integration Cheat Sheet — OWASP (2026-08-03)