Errors started climbing eight minutes after the deploy finished. Half the channel wants to roll back now. The other half wants five more minutes to understand what is happening, and someone mentions that the release included a migration. The decision in front of you is not “is the release definitely the cause?” It is whether undoing it is safe, and cheaper than being wrong.
Roll back when three things are true: the impact plausibly started with the release, the previous version can safely run against what the new version has already written, and no narrower lever covers the change. You do not need proof. A safe rollback is also the fastest test of whether the release is involved.
- Is it plausibly the release? Impact began in the deploy window, or only on new-version instances.
- Is rollback safe? No forward-only migration, new stored formats, or irreversible side effects.
- Is there a narrower lever? A feature flag or traffic shift beats a full rollback.
- If rollback is unsafe: use a flag, a targeted revert, or a small fix forward.
Teams lose most of the time in this decision to two things: waiting for certainty they do not need, and discovering halfway through a rollback that it is not safe. This guide covers both, with a decision tree, a safety checklist, and the options besides a full rollback.
When should you roll back a release?
Roll back a release when user impact plausibly started with it, the previous version can run safely on the current data, and no faster or narrower mitigation is available. A rollback here means redeploying the last known good version of the service. It restores the old code. It does not restore the old data, and that difference drives most of the decision.
The order of the questions matters. Checking safety before debating cause saves time, because an unsafe rollback is off the table no matter how strong the evidence is.
Two things about the first question. “Started with the deploy” means the first minute the key signal moved, taken from metrics, not the time of the first alert or customer report. And a rolling or canary deploy gives you a second, stronger signal for free: if errors come only from instances running the new version, the release is involved regardless of timing.
How much evidence do you need before rolling back?
You need enough evidence to make the release a plausible cause, not enough to explain the incident. The costs are asymmetric. If the rollback is safe and the release was innocent, you lose a few minutes and a deploy, and you have ruled out the most likely suspect. If you wait for proof and the release was guilty, users stay affected for the whole investigation.
Treat a safe rollback as an experiment. Recovery on the affected signal confirms the release was involved, which is exactly the evidence the investigation was trying to collect. No recovery rules it out in minutes and points you elsewhere.
Some signals make the release more or less likely. None of them has to be conclusive:
| Signal | What it suggests |
|---|---|
| Errors only on instances running the new version | Strongly points to the release. The best signal a canary or rolling deploy gives you. |
| First bad minute matches the deploy start or a traffic shift | Points to the release, unless a flag or config change happened in the same minute. |
| Stack traces in files or endpoints the release changed | Points to the release, and often to the specific change. |
| Impact started before the deploy | Points away. Look at flags, config, dependencies, and traffic. |
| Instances not yet updated fail the same way | Points away from the code, toward something shared: database, dependency, or config. |
| A dependency reports its own incident at the same time | Ambiguous. The release may have changed how you call it, or it may be unrelated. |
Lining these up is faster when logs, traces, and deploy markers share one timeline; How to Correlate Logs, Traces, and a Code Change During an Incident shows how. For timeouts specifically, API Timeouts After Deployment: A Practical Triage Checklist covers the first ten minutes.
If you use SLOs, the error budget gives the “is the impact material” question a number. The Google SRE workbook chapter on alerting on SLOs works through burn rates: at a 14.4× burn rate, one hour consumes 2% of a 30-day budget. A fast burn that started with a deploy is a clear case for rolling back first.
What makes a rollback unsafe?
A rollback is unsafe when the new version changed something the old version cannot handle and the rollback will not change back. Redeploying old code does not undo anything outside the code: the schema, the rows written, the messages queued, the cache entries stored, the emails sent. The old version wakes up in a world that the new version has already modified.
| Hazard | Why the old version breaks | What to do instead |
|---|---|---|
| Column dropped or renamed | Old queries reference a column that no longer exists | Fix forward, or restore the column first; plan expand and contract next time |
| New enum or status values stored | Old code cannot parse rows or messages it has never seen | Roll back only if the old version tolerates unknown values; otherwise flag off the writer |
| New message or event format | Old consumers fail on messages already queued | Drain or transform the queue, or keep new consumers running |
| Cache entries in a new shape or key | Old code misreads them, or leftovers go stale | Evict the affected keys as part of the rollback |
| External side effects | Emails, payments, and webhooks are not undone by a deploy | Roll back the code, then reconcile the side effects separately |
| Other services already upgraded | They call an endpoint or field the old version lacks | Roll back in dependency order, or keep the new API and flag off the behavior |
| Previous artifact not available | A rebuild from the old tag pulls different dependency versions | Keep immutable, versioned artifacts; redeploy by digest, not by rebuilding |
The most common hazard is the migration. A release that renames a column in one step is not safe to roll back once the migration has run:
ALTER TABLE users RENAME COLUMN email TO email_address;
-- After rolling back to v41, every query that reads users.email fails:
ERROR: column "email" does not exist
-- v42 (expand): add the new column; code writes both, still reads the old one
ALTER TABLE users ADD COLUMN email_address text;
-- Backfill in batches, then v43 reads email_address (still writes both)
UPDATE users SET email_address = email
WHERE email_address IS NULL AND id BETWEEN 1 AND 10000;
-- v44 (contract), once no version that reads users.email can be deployed:
ALTER TABLE users DROP COLUMN email;
The same reasoning applies to anything the new version persists. Stale Data After a Release describes the cache version of this problem: entries written under a new key survive the rollback and come back when the release is redeployed.
Rollback, feature flag, traffic shift, or fix forward: which one?
A full rollback is one of several levers. The best one is the narrowest lever that removes the impact, because it changes the least and usually acts fastest.
| Option | Choose it when | Watch out for | What it involves |
|---|---|---|---|
| Feature flag off | The change is behind a flag and the old path still exists | Flags that also gate data writes; the old path may not handle new data | A config change; no build or deploy |
| Traffic shift back | Blue-green or canary, with the old version still running | Old instances scaled down; sessions or caches tied to the new version | A router or load balancer change |
| Redeploy previous artifact | The whole release is suspect and the previous build is known good | Everything in the “does not revert” list above | A normal rollout of an existing artifact |
| Revert one change | Evidence points at one PR and the rest of the release is healthy | Revert conflicts; the revert must go through CI | Build, CI, and a rollout |
| Fix forward | Rollback is unsafe, or the fix is tiny and well understood | A rushed fix is new, lightly tested code shipped under pressure | Write, review, build, CI, and a rollout |
The mechanics of reverting a single commit (including git revert -m 1 for merge commits and what happens when you re-merge) are covered in Which Commit Caused the Production Incident?. The main difference between the options during an incident is how many steps stand between the decision and recovery:
The question is not “are we sure it’s the release?” It is “is undoing the release safe, and cheaper than being wrong?”
How to make the rollback decision during an incident, step by step
- Establish the impact and when it started. Which users and endpoints are affected, and the first minute the key signal moved, from metrics rather than from the first alert.
- Line it up against every change. Deploys, feature flag changes, configuration changes, and dependency incidents. Check whether errors come only from instances running the new version.
- Check rollback safety. Can the previous version run against the schema, data, messages, and cache entries the new version has written? Do any side effects need separate handling?
- Pick the narrowest safe lever. A flag or traffic shift if one covers the change, then a full rollback, then a targeted revert or fix forward when rollback is unsafe.
- Announce the decision and act. State the decision, the reason, and the signal you expect to recover in the incident channel, then execute it.
- Verify recovery on the same signal. Watch the signal that showed the impact return to baseline. A finished deploy is not a recovered service.
- Preserve evidence and block the bad version. Keep the bad SHA, logs, and traces. Stop the pipeline from redeploying the rolled-back version on the next merge.
- Plan the way back. Find the cause, add the missing check, and re-release with a smaller or guarded change.
On Kubernetes, the rollback itself is short. Know which revision you are going back to before you run it:
# Which revisions exist, and which one was last known good?
kubectl rollout history deployment/checkout
# Go back one revision, or to a specific one
kubectl rollout undo deployment/checkout
kubectl rollout undo deployment/checkout --to-revision=41
# Wait for it to finish, then watch the impact signal, not just this
kubectl rollout status deployment/checkout
# Helm-managed services: list releases, then roll back to a revision
helm history checkout
helm rollback checkout 41
kubectl rollout undo restores the Deployment’s previous pod template. It does not restore the contents of ConfigMaps or Secrets the pods read, it does not touch the database, and it does not stop your CD system from reapplying the new version on its next sync. If you deploy through GitOps, roll back in Git or pause the sync, or the rollback will be undone for you. See the Kubernetes documentation on rolling back a Deployment.
Decide in advance who may call a rollback. The Google SRE book chapter on managing incidents describes an incident commander who holds overall responsibility for the response and assigns the work, which is where a call like this belongs. A team that pre-authorizes rollback for releases marked safe decides in minutes. A team that needs the original author’s agreement debates while users wait.
How do you make a release safe to roll back?
Rollback safety is decided when the change is written, not during the incident. By the time errors are climbing, a one-step column rename or a new stored enum value has already made the choice for you. The checks belong in review:
- Does every migration keep the previous version working?
- Can the previous version read every value this one writes?
- Do message, event, and cache formats stay readable by it?
- What does this change send or charge that a rollback cannot undo?
- Is there a reconciliation path if it is rolled back mid-flight?
- Is risky behavior behind a flag with a working old path?
- Is the previous artifact kept and deployable by digest?
- Is the rollback owner and procedure written down?
These overlap with the deploy-time escapes in A PR Passed CI but Broke Production: What Did the Tests Miss?: a change that is unsafe to roll back is usually also a change that fails during a rolling deploy. The production readiness checklist includes rollback as a standing item for new services.
What should happen after the rollback?
A rollback ends the impact, not the incident. Three things keep it from happening again, or from being undone by accident:
- Freeze the bad version. Make sure the next merge to main does not redeploy it. Pause auto-deploy, or revert on the main branch so the pipeline builds a good version.
- Attribute the cause. The rollback tells you the release was involved, not which change. Narrow it down with the commit attribution steps, then write it up so the trigger, code path, and user impact connect, as in How to Write an RCA That Connects the Trigger, Code Path, and User Impact.
- Re-release smaller. Ship the good parts of the release without the bad change, and ship the fixed change behind a flag or a canary with automatic rollback.
How Tomosu helps
Most rollback hazards are visible in the pull request that creates them, long before anyone needs to roll back. Tomosu analyzes each change against the rest of the repository and surfaces the ones that affect rollback safety:
- Forward-only migrations: column drops, renames, and type changes that the previous version’s queries still depend on.
- New persisted formats: new enum values, message shapes, and cache keys that older readers may not handle.
- Side effects: new code paths that send, charge, or call out, where a rollback would need reconciliation.
- Blast radius: which endpoints, jobs, and services depend on the change, which shapes both the rollback decision and its order.
These feed the Production Reliability Index, so a release that will be hard to undo is known before it ships, and the on-call engineer is not the first to find out.
Scan your repository with Tomosu →
Key takeaways
- Roll back when impact plausibly started with the release, the old version can run on the current data, and no narrower lever covers the change.
- You do not need proof. A safe rollback is the fastest experiment for whether the release is involved.
- Check safety before cause. A rollback restores code, not schema, data, messages, caches, or side effects.
- Prefer the narrowest lever: flag off, then traffic shift, then full rollback, then revert or fix forward.
- Verify recovery on the signal that showed the impact, and block the bad version from redeploying.
- Rollback safety is decided in review: expand and contract, tolerant readers, and flags make every release reversible.
Frequently asked questions
When should you roll back a release?
Roll back when user impact started with the deploy, the previous version can safely run against the data the new version has written, and no narrower lever such as a feature flag covers the change. You do not need proof that the release caused the problem. If rollback is safe and cheap, it is also the fastest way to find out.
Should I roll back or roll forward?
Roll back by default when it is safe, because redeploying a known good artifact is usually faster and less risky than writing a fix under pressure. Roll forward when rollback is unsafe, for example after a forward-only migration or a data format change, or when the fix is tiny, well understood, and can ship faster than the rollback.
How much evidence do I need before rolling back?
Less than you need to explain the incident. If impact started within the deploy window or appears only on instances running the new version, and rollback is safe, that is enough. Rolling back is a reversible experiment: recovery confirms the release was involved, and no recovery rules it out quickly.
What makes a rollback unsafe?
Anything the new version changed that the old version cannot handle: a dropped or renamed column, new enum or status values in stored rows, messages or cache entries in a new format, other services that already depend on a new API, and side effects such as emails or payments that a code rollback does not undo.
Does rolling back a deployment undo database migrations?
No. Redeploying the previous version changes the running code, not the database. Schema changes and rows written by the new version stay in place. That is why migrations should be backward compatible with the previous version, using expand and contract, so a code rollback stays safe.
What should I do after rolling back a release?
Confirm recovery on the signal that showed the impact, keep the bad version’s SHA, logs, and traces, and block the pipeline from redeploying it automatically. Then find the cause, add the check that would have caught it, and re-release with a smaller or feature-flagged change.
Who should decide to roll back during an incident?
The incident commander or on-call engineer, with authority agreed before the incident. Rollback should not require consensus or the original author’s approval. Teams that pre-authorize rollback for safe releases decide in minutes instead of debating while users are affected.
Whether a rollback is safe is decided when the change is written, not during the incident. Tomosu flags forward-only migrations and new persisted formats in the pull request, while they are still cheap to change. Assess your repository →