Incident history
This is the example page's own history. Everything below is what a subscriber received by email, in the order it was posted.
Webhook delivery delayed up to 40 minutes
Resolved. The backlog cleared at 17:44 UTC and delivery latency is back under 5 seconds. No webhooks were lost; all delayed events were delivered with their original timestamps.17:52 UTC
Monitoring. Queue depth is falling steadily. We expect the backlog to clear within the hour.16:30 UTC
Identified. A single delivery worker was retrying against a customer endpoint that accepted the connection and never responded, holding the slot for the full timeout. Timeout lowered from 120 s to 20 s and the worker pool widened.15:20 UTC
Investigating. Webhook delivery is running behind. Checkout API and Dashboard are unaffected.14:40 UTC
Elevated latency on Checkout API
Resolved. p99 is back to 180 ms. Cause was a slow query introduced in the morning deploy; the index has been added and the deploy is unblocked.09:31 UTC
Investigating. p99 latency on the Checkout API rose from 180 ms to 2.1 s at 09:05 UTC. Requests are succeeding, just slowly.09:08 UTC
Scheduled maintenance — settlement database
Completed. Failover finished ahead of the window. The nightly batch ran normally the same evening.03:05 UTC
Scheduled. Primary database failover between 02:00 and 04:00 UTC. Settlement batch may start up to an hour late; the API is unaffected.posted 29 Jul
What we ask of ourselves
Post the first update within fifteen minutes of noticing, even when the update is only "we see it and we do not know why yet". A status page that goes quiet during an incident is worse than no status page at all.