Leadership · Delivery

Release risk usually isn't in the code

When a release goes wrong, the first instinct is to look for the bug. Sometimes there is one. But most of the release risk I’ve dealt with had nothing to do with the code. The code did exactly what it was written to do. What caught us out was everything around it.

Where the risk actually lives

A few patterns keep coming back:

  • Environments that don’t match. An infrastructure upgrade that timed out in a lower environment and needed a restart before it validated. A pre-production environment whose state, configuration or content didn’t really represent production. The release passed its tests, just not against the thing it was going to.
  • Manual steps nobody owns. When a change needs manual packaging or installation, the release lead shouldn’t have to work out the steps from a ticket. The team delivering the change should hand over exact instructions. That sounds obvious until you’re the one doing it at night.
  • Late changes. Changes that land after the release branch is cut cause redeployments and fresh regression risk, and they’re almost always “small”.
  • Rollback in theory. “We can roll back” is only a plan if someone has worked out what rollback means, including data, and checked it’s possible.

None of these show up in a code review, and none of them is anyone’s job by default.

What preparing a migration taught me

The clearest example was a migration of a distributed data store that customer sessions depend on. On paper, it was a connection string change. In practice, it was a list of questions nobody had asked yet.

I wrote the plan around how it could fail, not just how it should work:

  • Test where it’s safe first. Lower-environment testing, connectivity checks and regression validation before anything touched production.
  • Keep the old platform alive. The previous store stayed in place during validation, so rollback meant restoring the previous connection and application version, not rebuilding anything.
  • Don’t decommission early. The old platform was only retired once the new one had proven itself, not on the day of cutover.
  • Think about the customers mid-journey. A straight cutover would have dropped active sessions. I proposed running both stores side by side for a while: new writes to the new store, with the old one as a read fallback. In the end the team chose a copy-and-swap approach instead, which was a fair call, but the point was making session continuity an explicit decision rather than a surprise.
  • Write down the risk. What the change is for, what it affects, what happens if we don’t do it, how it could fail and how we’d mitigate each one.

During the work, a compatibility problem with the new configuration made one operation fail while everything else looked healthy. The temptation was to fold the fix into normal development. I asked for it to be isolated and shipped in its own release instead, so it could be tested and rolled back on its own.

Feature complete isn’t production ready

Teams are measured on finishing features, so “done” naturally means “it works”. A release needs a second definition of done: we’re confident it will behave in production, and we know what to do if it doesn’t. Before anything significant goes out, I ask:

  1. What’s different between where we tested and where this is going?
  2. Who owns each step on the day, by name, including the manual ones?
  3. What changed after the branch was cut, and why?
  4. How will we know it worked, and what would tell us it hasn’t?
  5. What’s the rollback, and what happens to the data?

Don’t be the only one who knows

The last lesson is about people. If release knowledge sits with one person, every release depends on them being available, and the lead becomes the bottleneck. So I’ve started handing release leadership to other engineers, with an experienced engineer shadowing them through each step. The first time takes longer. After that, the team has more than one person who can run a release safely, which is the actual goal.