Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Deployment

How to Decide Whether to Roll Back a Release

Tomosu AI·14 min read·

Errors started climbing eight minutes after the deploy finished. Half the channel wants to roll back now. The other half wants five more minutes to understand what is happening, and someone mentions that the release included a migration. The decision in front of you is not “is the release definitely the cause?” It is whether undoing it is safe, and cheaper than being wrong.

Quick answer

Roll back when three things are true: the impact plausibly started with the release, the previous version can safely run against what the new version has already written, and no narrower lever covers the change. You do not need proof. A safe rollback is also the fastest test of whether the release is involved.

Teams lose most of the time in this decision to two things: waiting for certainty they do not need, and discovering halfway through a rollback that it is not safe. This guide covers both, with a decision tree, a safety checklist, and the options besides a full rollback.

When should you roll back a release?

Roll back a release when user impact plausibly started with it, the previous version can run safely on the current data, and no faster or narrower mitigation is available. A rollback here means redeploying the last known good version of the service. It restores the old code. It does not restore the old data, and that difference drives most of the decision.

The order of the questions matters. Checking safety before debating cause saves time, because an unsafe rollback is off the table no matter how strong the evidence is.

SHOULD WE ROLL BACK? FOUR QUESTIONS, IN ORDER Q1 · TIMING AND VERSION Did impact start with the deploy, or only on new pods? Q2 · ROLLBACK SAFETY Can the old version run on data the new one wrote? Q3 · NARROWER LEVER Is a feature flag or traffic shift ready for this change? Q4 · IMPACT Is the impact ongoing and material? Look beyond this release Flags, config, dependencies Rollback is unsafe Flag, targeted revert, fix forward Use the narrower lever Keeps the rest of the release Roll back now Diagnose after users recover NONOYESYES YESYESNONO Time-box the investigation Set a deadline now. If the cause is not explained by then, roll back. Minor impact is not a reason to wait forever.
Safety comes before certainty. If the old version cannot run on the new data, stronger evidence does not make a rollback safer.

Two things about the first question. “Started with the deploy” means the first minute the key signal moved, taken from metrics, not the time of the first alert or customer report. And a rolling or canary deploy gives you a second, stronger signal for free: if errors come only from instances running the new version, the release is involved regardless of timing.

How much evidence do you need before rolling back?

You need enough evidence to make the release a plausible cause, not enough to explain the incident. The costs are asymmetric. If the rollback is safe and the release was innocent, you lose a few minutes and a deploy, and you have ruled out the most likely suspect. If you wait for proof and the release was guilty, users stay affected for the whole investigation.

Treat a safe rollback as an experiment. Recovery on the affected signal confirms the release was involved, which is exactly the evidence the investigation was trying to collect. No recovery rules it out in minutes and points you elsewhere.

SAFETY MATTERS MORE THAN CERTAINTY ROLLBACK SAFE ROLLBACK UNSAFE Roll back anyway A cheap experiment. Recovery confirms the release; no recovery rules it out in minutes. Roll back now The default. Diagnose after users recover, with the bad version’s evidence preserved. Mitigate without deploying Scale out, shed load, disable a feature, fail over. Keep investigating. Narrow lever or fix forward Flag off, revert one change, or ship a minimal fix. Keep it small: it is new code under pressure. LOW CONFIDENCE HIGH CONFIDENCE THE RELEASE CAUSED IT Both top quadrants say roll back. Confidence changes the action only when rollback is unsafe.
When rollback is safe, the answer is the same at any confidence level. The hard cases are all in the bottom row.

Some signals make the release more or less likely. None of them has to be conclusive:

SignalWhat it suggests
Errors only on instances running the new versionStrongly points to the release. The best signal a canary or rolling deploy gives you.
First bad minute matches the deploy start or a traffic shiftPoints to the release, unless a flag or config change happened in the same minute.
Stack traces in files or endpoints the release changedPoints to the release, and often to the specific change.
Impact started before the deployPoints away. Look at flags, config, dependencies, and traffic.
Instances not yet updated fail the same wayPoints away from the code, toward something shared: database, dependency, or config.
A dependency reports its own incident at the same timeAmbiguous. The release may have changed how you call it, or it may be unrelated.

Lining these up is faster when logs, traces, and deploy markers share one timeline; How to Correlate Logs, Traces, and a Code Change During an Incident shows how. For timeouts specifically, API Timeouts After Deployment: A Practical Triage Checklist covers the first ten minutes.

If you use SLOs, the error budget gives the “is the impact material” question a number. The Google SRE workbook chapter on alerting on SLOs works through burn rates: at a 14.4× burn rate, one hour consumes 2% of a 30-day budget. A fast burn that started with a deploy is a clear case for rolling back first.

What makes a rollback unsafe?

A rollback is unsafe when the new version changed something the old version cannot handle and the rollback will not change back. Redeploying old code does not undo anything outside the code: the schema, the rows written, the messages queued, the cache entries stored, the emails sent. The old version wakes up in a world that the new version has already modified.

ROLL BACK FROM V42 TO V41: WHAT ACTUALLY CHANGES BACK? REVERTS WITH THE CODE Container image and code paths Settings in the deployment spec New endpoints and handlers In-process state (on restart) This is all a rollback promises. DOES NOT REVERT Schema changes from migrations Rows written with new values or formats Queued messages in a new format Cache entries under new keys or shapes Emails, webhooks, payments already sent ConfigMaps, secrets, flags changed apart Other services now using a new API Rollback safety is one question: can v41 run correctly against everything in the right-hand box?
The data outlives the code. A rollback is safe only if the previous version tolerates what the new version left behind.
HazardWhy the old version breaksWhat to do instead
Column dropped or renamedOld queries reference a column that no longer existsFix forward, or restore the column first; plan expand and contract next time
New enum or status values storedOld code cannot parse rows or messages it has never seenRoll back only if the old version tolerates unknown values; otherwise flag off the writer
New message or event formatOld consumers fail on messages already queuedDrain or transform the queue, or keep new consumers running
Cache entries in a new shape or keyOld code misreads them, or leftovers go staleEvict the affected keys as part of the rollback
External side effectsEmails, payments, and webhooks are not undone by a deployRoll back the code, then reconcile the side effects separately
Other services already upgradedThey call an endpoint or field the old version lacksRoll back in dependency order, or keep the new API and flag off the behavior
Previous artifact not availableA rebuild from the old tag pulls different dependency versionsKeep immutable, versioned artifacts; redeploy by digest, not by rebuilding

The most common hazard is the migration. A release that renames a column in one step is not safe to roll back once the migration has run:

Migration in v42rollback breaks v41
ALTER TABLE users RENAME COLUMN email TO email_address;

-- After rolling back to v41, every query that reads users.email fails:
ERROR:  column "email" does not exist
Expand and contracteach step safe to roll back
-- v42 (expand): add the new column; code writes both, still reads the old one
ALTER TABLE users ADD COLUMN email_address text;

-- Backfill in batches, then v43 reads email_address (still writes both)
UPDATE users SET email_address = email
WHERE email_address IS NULL AND id BETWEEN 1 AND 10000;

-- v44 (contract), once no version that reads users.email can be deployed:
ALTER TABLE users DROP COLUMN email;

The same reasoning applies to anything the new version persists. Stale Data After a Release describes the cache version of this problem: entries written under a new key survive the rollback and come back when the release is redeployed.

Rollback, feature flag, traffic shift, or fix forward: which one?

A full rollback is one of several levers. The best one is the narrowest lever that removes the impact, because it changes the least and usually acts fastest.

OptionChoose it whenWatch out forWhat it involves
Feature flag offThe change is behind a flag and the old path still existsFlags that also gate data writes; the old path may not handle new dataA config change; no build or deploy
Traffic shift backBlue-green or canary, with the old version still runningOld instances scaled down; sessions or caches tied to the new versionA router or load balancer change
Redeploy previous artifactThe whole release is suspect and the previous build is known goodEverything in the “does not revert” list aboveA normal rollout of an existing artifact
Revert one changeEvidence points at one PR and the rest of the release is healthyRevert conflicts; the revert must go through CIBuild, CI, and a rollout
Fix forwardRollback is unsafe, or the fix is tiny and well understoodA rushed fix is new, lightly tested code shipped under pressureWrite, review, build, CI, and a rollout

The mechanics of reverting a single commit (including git revert -m 1 for merge commits and what happens when you re-merge) are covered in Which Commit Caused the Production Incident?. The main difference between the options during an incident is how many steps stand between the decision and recovery:

STEPS BETWEEN THE DECISION AND RECOVERY (ILLUSTRATIVE) Feature flag off Traffic shift back Redeploy previous Revert one change Fix forward config routing rollout build + CI rollout write + review fix build + CI rollout TIME TO RECOVERY Every added step adds time and a new chance to fail. Measure these for your own pipeline before you need them.
A flag or traffic shift acts in one step. A fix forward adds writing, review, and CI to the same rollout a rollback needs.

The question is not “are we sure it’s the release?” It is “is undoing the release safe, and cheaper than being wrong?”

How to make the rollback decision during an incident, step by step

  1. Establish the impact and when it started. Which users and endpoints are affected, and the first minute the key signal moved, from metrics rather than from the first alert.
  2. Line it up against every change. Deploys, feature flag changes, configuration changes, and dependency incidents. Check whether errors come only from instances running the new version.
  3. Check rollback safety. Can the previous version run against the schema, data, messages, and cache entries the new version has written? Do any side effects need separate handling?
  4. Pick the narrowest safe lever. A flag or traffic shift if one covers the change, then a full rollback, then a targeted revert or fix forward when rollback is unsafe.
  5. Announce the decision and act. State the decision, the reason, and the signal you expect to recover in the incident channel, then execute it.
  6. Verify recovery on the same signal. Watch the signal that showed the impact return to baseline. A finished deploy is not a recovered service.
  7. Preserve evidence and block the bad version. Keep the bad SHA, logs, and traces. Stop the pipeline from redeploying the rolled-back version on the next merge.
  8. Plan the way back. Find the cause, add the missing check, and re-release with a smaller or guarded change.

On Kubernetes, the rollback itself is short. Know which revision you are going back to before you run it:

Shell · Kubernetes and Helm
# Which revisions exist, and which one was last known good?
kubectl rollout history deployment/checkout

# Go back one revision, or to a specific one
kubectl rollout undo deployment/checkout
kubectl rollout undo deployment/checkout --to-revision=41

# Wait for it to finish, then watch the impact signal, not just this
kubectl rollout status deployment/checkout

# Helm-managed services: list releases, then roll back to a revision
helm history checkout
helm rollback checkout 41
What kubectl rollout undo does not cover

kubectl rollout undo restores the Deployment’s previous pod template. It does not restore the contents of ConfigMaps or Secrets the pods read, it does not touch the database, and it does not stop your CD system from reapplying the new version on its next sync. If you deploy through GitOps, roll back in Git or pause the sync, or the rollback will be undone for you. See the Kubernetes documentation on rolling back a Deployment.

Decide in advance who may call a rollback. The Google SRE book chapter on managing incidents describes an incident commander who holds overall responsibility for the response and assigns the work, which is where a call like this belongs. A team that pre-authorizes rollback for releases marked safe decides in minutes. A team that needs the original author’s agreement debates while users wait.

How do you make a release safe to roll back?

Rollback safety is decided when the change is written, not during the incident. By the time errors are climbing, a one-step column rename or a new stored enum value has already made the choice for you. The checks belong in review:

Data
  • Does every migration keep the previous version working?
  • Can the previous version read every value this one writes?
  • Do message, event, and cache formats stay readable by it?
Side effects
  • What does this change send or charge that a rollback cannot undo?
  • Is there a reconciliation path if it is rolled back mid-flight?
Levers
  • Is risky behavior behind a flag with a working old path?
  • Is the previous artifact kept and deployable by digest?
  • Is the rollback owner and procedure written down?

These overlap with the deploy-time escapes in A PR Passed CI but Broke Production: What Did the Tests Miss?: a change that is unsafe to roll back is usually also a change that fails during a rolling deploy. The production readiness checklist includes rollback as a standing item for new services.

What should happen after the rollback?

A rollback ends the impact, not the incident. Three things keep it from happening again, or from being undone by accident:

How Tomosu helps

Most rollback hazards are visible in the pull request that creates them, long before anyone needs to roll back. Tomosu analyzes each change against the rest of the repository and surfaces the ones that affect rollback safety:

These feed the Production Reliability Index, so a release that will be hard to undo is known before it ships, and the on-call engineer is not the first to find out.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

When should you roll back a release?

Roll back when user impact started with the deploy, the previous version can safely run against the data the new version has written, and no narrower lever such as a feature flag covers the change. You do not need proof that the release caused the problem. If rollback is safe and cheap, it is also the fastest way to find out.

Should I roll back or roll forward?

Roll back by default when it is safe, because redeploying a known good artifact is usually faster and less risky than writing a fix under pressure. Roll forward when rollback is unsafe, for example after a forward-only migration or a data format change, or when the fix is tiny, well understood, and can ship faster than the rollback.

How much evidence do I need before rolling back?

Less than you need to explain the incident. If impact started within the deploy window or appears only on instances running the new version, and rollback is safe, that is enough. Rolling back is a reversible experiment: recovery confirms the release was involved, and no recovery rules it out quickly.

What makes a rollback unsafe?

Anything the new version changed that the old version cannot handle: a dropped or renamed column, new enum or status values in stored rows, messages or cache entries in a new format, other services that already depend on a new API, and side effects such as emails or payments that a code rollback does not undo.

Does rolling back a deployment undo database migrations?

No. Redeploying the previous version changes the running code, not the database. Schema changes and rows written by the new version stay in place. That is why migrations should be backward compatible with the previous version, using expand and contract, so a code rollback stays safe.

What should I do after rolling back a release?

Confirm recovery on the signal that showed the impact, keep the bad version’s SHA, logs, and traces, and block the pipeline from redeploying it automatically. Then find the cause, add the check that would have caught it, and re-release with a smaller or feature-flagged change.

Who should decide to roll back during an incident?

The incident commander or on-call engineer, with authority agreed before the incident. Rollback should not require consensus or the original author’s approval. Teams that pre-authorize rollback for safe releases decide in minutes instead of debating while users are affected.


Whether a rollback is safe is decided when the change is written, not during the incident. Tomosu flags forward-only migrations and new persisted formats in the pull request, while they are still cheap to change. Assess your repository →