The First Five Minutes of an Incident
The first five minutes of an incident are rarely long enough to diagnose it. They are long enough to damage the evidence, widen the uncertainty, and create three competing versions of what happened.
What I built, what broke, and the evidence that changed the next decision.
No spam. No marketing. Just the writing.
The first five minutes of an incident are rarely long enough to diagnose it. They are long enough to damage the evidence, widen the uncertainty, and create three competing versions of what happened.
The rollback used to be the last part of my change notes. I’d describe what I wanted to alter, list the commands, think through validation, and then add a reassuring line about putting things back if necessary.
Small systems have nowhere to hide their assumptions. That can feel limiting when I’m building one, but it’s a gift when I have to operate it later.
A restart is satisfying because it produces motion. The process stops, the process starts, and the terminal returns a clean status. That can be exactly the right repair. It can also leave the system holding the same bad state, an unprocessed queue, and several users who are still refreshing a bro...
CEOs say AI isn't moving the needle on productivity. I've been running AI in production for months. Here's why they're both right and wrong — and what the actual measurement problem is.