# Production Readiness Checklist for a Backend Service

Source: https://tomosu.ai/blogs/production-readiness-checklist-backend-service.html  
License: Free to copy and adapt. Attribution appreciated, not required.  
See also: Mercari's public checklist (MIT), https://github.com/mercari/production-readiness-checklist  
Version: 2026-09-29 - 51 items in 9 sections

How to use: copy this file into the service repository (for example docs/production-readiness.md) or into a launch ticket. Check an item only when you can link evidence: a dashboard, an alert rule, a test run, a runbook, or a pull request. Mark items that do not apply as N/A with a one-line reason. Re-run the checklist when the service changes tier, owner, or architecture.

## Ownership & on-call

- [ ] A named owning team is recorded in the service catalog  
  _Why:_ Incidents need someone to page and decisions need someone to make them. _Owner:_ Engineering manager
- [ ] The service has a criticality tier that sets how strict this checklist is  
  _Why:_ A tier-1 checkout API and an internal batch job should not face the same bar. _Owner:_ Tech lead
- [ ] An on-call rotation of at least two people is live and has received a test page  
  _Why:_ An alert that reaches nobody is not an alert. _Owner:_ Engineering manager
- [ ] Escalation contacts for every critical dependency are documented  
  _Why:_ Many incidents are resolved by another team; finding them should not take an hour. _Owner:_ On-call
- [ ] On-call engineers know the incident process and severity levels  
  _Why:_ Declaring, communicating, and handing off an incident is a skill, not a document. _Owner:_ On-call

## Observability

- [ ] SLIs are defined for availability and latency of user-facing operations  
  _Why:_ Without an SLI there is no shared definition of working. _Owner:_ Tech lead
- [ ] SLOs and an error budget are agreed, with a documented response when the budget is spent  
  _Why:_ An SLO nobody acts on is a dashboard, not a commitment. _Owner:_ Tech lead + product
- [ ] Paging alerts fire on SLO burn rate or user-visible symptoms, not on CPU or single errors  
  _Why:_ Cause-based alerts page for things users never notice and miss things they do. _Owner:_ SRE / on-call
- [ ] Every paging alert links to a runbook  
  _Why:_ The engineer paged at 3 a.m. needs the next step, not a graph. _Owner:_ On-call
- [ ] A dashboard shows traffic, errors, latency, and saturation per endpoint and per dependency  
  _Why:_ The four golden signals localize most incidents in minutes. _Owner:_ Service team
- [ ] Logs are structured and every line carries a request or trace id  
  _Why:_ Correlation ids turn thousands of lines into one request's story. _Owner:_ Service team
- [ ] Logs contain no secrets, tokens, or unnecessary personal data  
  _Why:_ Logs are copied widely and retained long; leaks there are hard to clean up. _Owner:_ Service team + security
- [ ] Trace context (W3C traceparent) propagates across inbound calls, outbound calls, and messages  
  _Why:_ A trace that stops at the queue cannot explain an async failure. _Owner:_ Service team
- [ ] Metrics carry the deployed version so a change can be correlated with a regression  
  _Why:_ The first incident question is usually what changed. _Owner:_ Service team

## Reliability

- [ ] Every outbound call (HTTP, gRPC, database, cache, queue) has an explicit timeout  
  _Why:_ Default timeouts are often infinite or far longer than your own deadline. _Owner:_ Service team
- [ ] Timeouts fit inside the caller's deadline, leaving room for a retry if one is allowed  
  _Why:_ A downstream timeout longer than the upstream one wastes work nobody waits for. _Owner:_ Service team
- [ ] Retries are capped, use exponential backoff with jitter, and only wrap idempotent operations  
  _Why:_ Unbounded synchronized retries turn a blip into a retry storm. _Owner:_ Service team
- [ ] Retries happen at one layer only  
  _Why:_ Three layers retrying three times each send 27 attempts to a struggling dependency. _Owner:_ Tech lead
- [ ] Write endpoints that clients may retry are idempotent or accept an idempotency key  
  _Why:_ Timeouts and retries guarantee duplicates; idempotency makes them harmless. _Owner:_ Service team
- [ ] On SIGTERM the service stops taking new work, drains in-flight requests, and exits within the grace period  
  _Why:_ Without graceful shutdown every deploy drops requests. _Owner:_ Service team
- [ ] The readiness probe reflects ability to serve; the liveness probe checks only the process  
  _Why:_ A liveness probe that checks the database restarts every pod when the database blips. _Owner:_ Service team
- [ ] CPU and memory requests and limits are set from measured usage, and runtime memory fits inside the limit  
  _Why:_ An unset limit starves neighbors; a wrong one causes OOMKilled restarts. _Owner:_ Service team + platform
- [ ] At least two replicas run across zones, with autoscaling bounds and a PodDisruptionBudget  
  _Why:_ One replica means every node drain or deploy is an outage. _Owner:_ Platform

## Dependencies

- [ ] Every dependency is listed with its owner, SLO, and whether it is hard or soft  
  _Why:_ Your availability cannot exceed that of your hard dependencies combined. _Owner:_ Tech lead
- [ ] Owners of shared dependencies have confirmed capacity for your expected peak  
  _Why:_ A new service can be the load test a shared database never had. _Owner:_ Tech lead
- [ ] Behavior is known and tested when each dependency is slow, down, or returning errors  
  _Why:_ Slow is worse than down: it holds threads, connections, and memory. _Owner:_ Service team
- [ ] Soft dependencies have a fallback (cached value, default, or degraded response)  
  _Why:_ A recommendation widget should not take down checkout. _Owner:_ Service team
- [ ] Concurrency to each dependency is bounded by pools, bulkheads, or a circuit breaker  
  _Why:_ One slow dependency should not consume every worker in the process. _Owner:_ Service team

## Data

- [ ] Schema migrations are backward compatible with the previously deployed version  
  _Why:_ During a rollout and after a rollback, old and new code share one schema. _Owner:_ Service team
- [ ] Migrations have been run against production-sized data to measure lock time and duration  
  _Why:_ A migration that takes milliseconds on staging can lock a large table for minutes. _Owner:_ Service team + DBA
- [ ] Backups are automated and meet a stated recovery point objective (RPO)  
  _Why:_ Without a stated RPO nobody knows how much data loss is acceptable. _Owner:_ Platform / DBA
- [ ] A restore has been tested end to end within the recovery time objective (RTO), and the date is recorded  
  _Why:_ A backup that has never been restored is a hope, not a backup. _Owner:_ Platform / DBA
- [ ] Data is classified, and sensitive data is encrypted at rest with audited access  
  _Why:_ Classification drives retention, access, and breach obligations. _Owner:_ Service team + security

## Security

- [ ] Secrets live in a secret manager, not in code, images, or committed config, and rotation has been tested  
  _Why:_ A secret you cannot rotate quickly is a standing incident. _Owner:_ Service team + security
- [ ] Every endpoint requires authentication, and authorization is checked per resource  
  _Why:_ Missing object-level checks let one user read another user's data. _Owner:_ Service team
- [ ] The service runs with a least-privilege identity and restricted network access  
  _Why:_ Limits what an attacker, or a bug, can reach. _Owner:_ Platform + security
- [ ] Dependency and container image scanning runs in CI with a policy for critical findings  
  _Why:_ Known vulnerabilities in dependencies are among the most common ways in. _Owner:_ Service team
- [ ] Inputs are validated, request sizes are limited, and traffic is encrypted in transit  
  _Why:_ Unbounded input is both a security and an availability problem. _Owner:_ Service team

## Deployment

- [ ] Deploys run from CI/CD with no manual steps on production hosts  
  _Why:_ Manual steps are skipped, reordered, and unrecorded under pressure. _Owner:_ Service team
- [ ] Rollout is progressive (canary or percentage) with automated checks on SLIs between stages  
  _Why:_ A bad change should reach 1% of traffic, not 100%. _Owner:_ Service team + platform
- [ ] Risky behavior ships behind a feature flag that can be turned off without a deploy  
  _Why:_ Turning off a flag takes seconds; a rollback pipeline can take much longer. _Owner:_ Service team
- [ ] Rollback has been exercised on this service and its duration is known  
  _Why:_ An untested rollback fails exactly when you need it. _Owner:_ Service team
- [ ] Configuration changes get the same review and progressive rollout as code  
  _Why:_ Many outages start as a one-line config change. _Owner:_ Tech lead

## Capacity

- [ ] Expected peak traffic is estimated, including growth and planned events  
  _Why:_ Capacity work needs a number to aim at. _Owner:_ Tech lead + product
- [ ] A load test at or above expected peak ran on a production-like environment  
  _Why:_ Queues, pools, and locks only show their limits under load. _Owner:_ Service team
- [ ] The saturation point and the first resource to run out are known  
  _Why:_ Knowing what breaks first tells you what to watch and what to scale. _Owner:_ Service team
- [ ] Inbound rate limits exist, and relevant cloud and third-party quotas are known  
  _Why:_ A quota you did not know about is an outage with no code change. _Owner:_ Service team + platform

## Documentation

- [ ] A runbook exists for every paging alert: what it means, how to confirm, how to mitigate  
  _Why:_ Runbooks turn one expert's knowledge into the whole rotation's. _Owner:_ On-call
- [ ] An architecture diagram and dependency list are current  
  _Why:_ Responders need to see the blast radius before they can contain it. _Owner:_ Tech lead
- [ ] The API contract is documented and versioned  
  _Why:_ Clients break silently when contracts change implicitly. _Owner:_ Service team
- [ ] Manual operations (replay, backfill, reprocess, data fix) are documented and safe to run twice  
  _Why:_ Recovery steps are run under pressure, often more than once. _Owner:_ Service team

---

Sign-off

- Service: 
- Tier: 
- Reviewed by: 
- Date: 
- Open items and owners: 
