$ git log --grep="incidents"
Incidents
Things that broke, and what they taught.
6 posts
- 5 min read
Your Retry Budget Is Longer Than Your Users' Patience
An internal dashboard failed intermittently. Five plausible causes measured and excluded, one measurement trap of my own making, and a database idle 180 hours out of 182. The real problem was a retry policy allowed to outlive the person waiting for it.
- 3 min read
Who Checks the Outside Check?
A reader asked how outside checks stay reliable over time. The honest answer is partly: what exists, a correction to last week's post, and the gap the question found.
- 5 min read
The Fix Took Production Down for Three Days
In August I wrote about a webhook that let broken services start, and I fixed it. Last week that fix refused to admit its own replacement. There is no correct setting here, only a choice of failure and a bill that comes with it.
- 3 min read
Your Alert Cleared Itself
Seven alerts fired when production stopped working. Two hours and forty-five minutes later the monitoring announced that telemetry had been restored. Nothing had been restored, and the cluster stayed down for another three days.
- 7 min read
What It Takes to Run the Door at a Real Event
Feria Arequipa trusted my platform with access control: real people, real doors, no second chances. The numbers, the Friday evening it broke during the event, and the three other things that went wrong that nobody noticed.
- 5 min read
49 Microservices, One Invisible Failure: When “Running” Doesn't Mean Working
The overengineered side project is now a 49-service platform running live events. A Friday-evening incident where every dashboard was green and nothing worked — and the one-line lesson it left.