A release has failed in production two hours after go-live. Walk me through what you do.
Suggested answer
I work the incident before I work the cause.
1. Establish the blast radius — which users, which processes, is data being written incorrectly. Data corruption changes the decision, because it makes waiting expensive.
2. Stabilise — if there is a safe way to reduce exposure immediately, such as switching off a feature flag or deactivating a flow, take it while the decision is being made.
3. Take the recovery decision against the plan we prepared, not on the spot: forward fix or redeploy the previous version, per component, with the elapsed time of each route known in advance. Two hours in, with records already created against the new configuration, a forward fix is often faster and safer than a revert, because reverting metadata does not un-create data.
4. Communicate — the named authoriser makes the call, and the business is told what is happening and what to do in the meantime.
Afterwards, the review. Why did staging not catch it — a data fidelity gap, a missing manual step, an untested integration? And whichever it was, the fix goes into the process, not into a person's memory.
The point I would make to any panel is that all of this depends on the recovery plan having been written and rehearsed before go-live. Improvising a rollback at 4pm on release day is how a bad release becomes a bad week.
Practice content for interview preparation; not an official vendor answer. Verify details against current product documentation.
Community comments (0)
No comments yet.
Sign in or create a free account to add a comment. Comments are moderated before they appear.