A production readiness checklist is the list of things a backend service must prove before it takes real traffic: someone owns it, you can see it failing, it survives its dependencies failing, its data can be restored, and a bad deploy can be undone. This one has 51 checkable items in nine sections, and you can download it, adapt it, and put it in your repository.
Take the checklist with you. Markdown with - [ ] boxes for a repo or pull request, CSV with owner and status columns for a spreadsheet or tracker. Free to copy and adapt.
A production readiness checklist for a backend service verifies, with evidence, that the service can be operated safely before launch. It is not a code review. It asks whether the service is owned, observable, resilient to its dependencies, recoverable, secure, and safe to deploy and roll back. Scale the bar to the service’s criticality.
- Ownership: a named team, a live on-call rotation, escalation paths.
- Observability: SLIs and SLOs, burn-rate alerts with runbooks, logs and traces joined by a correlation id.
- Reliability: timeouts, bounded retries with backoff and jitter, idempotency, graceful shutdown, correct probes, resource limits.
- Dependencies and data: known failure modes and fallbacks, backward-compatible migrations, a tested restore.
- Security, deployment, capacity, docs: managed secrets, progressive rollout with a tested rollback, a load test, runbooks.
The list below is opinionated but not exotic. It follows the same shape as public checklists such as Mercari’s production readiness checklist and the launch and review practices described in the Google SRE book. What this guide adds is the reasoning behind each item and the evidence to ask for, so the checklist is a review and not a ritual.
What is a production readiness checklist?
A production readiness checklist is a structured list of operational requirements a service must meet before it serves customers. Teams run it as part of a production readiness review (PRR): the service owners gather evidence for each item, someone outside the team reviews it, and gaps become tracked work with owners and dates.
The checklist exists because code that passes tests is not the same as a service that can be operated. Tests tell you the happy path works. The checklist asks what happens at 3 a.m. when the payment provider is slow, a node is drained mid-request, or last week’s migration needs to be rolled back. Most of those answers live in configuration, client code, and runbooks that no unit test exercises.
Two practical rules keep the checklist useful. First, scale it to the service’s criticality. Mercari’s checklist ties the required items to a production readiness level derived from the service’s SLO, so an internal tool is not held to the bar of a payment path. Second, every checked item needs evidence: a link to a dashboard, an alert rule, a load test result, a runbook, or a pull request. A checkbox without a link is an opinion.
Service launch checklist vs pre-deployment checklist: what is the difference?
People search for both and often mean different things. A service launch checklist is the one-time production readiness review for a new service or a major change in how it is used. A pre-deployment checklist for a backend service is the short set of checks every release passes, most of which should be automated in CI/CD.
| Checklist | When it runs | Typical items |
|---|---|---|
| Service launch checklist | Before first production traffic; again on tier, owner, or architecture change | Owner and on-call, SLOs, dashboards, dependency review, backup restore test, load test, runbooks |
| Pre-deployment checklist | Every release | Tests and scans pass, migration is backward compatible, new calls have timeouts, risky paths behind flags, rollout plan and rollback path known |
| Continuous checks | Always | SLO burn alerts, dependency and config drift, expiring certificates and secrets, restore tests on a schedule |
The rest of this guide is the launch checklist, because it is the superset. The pre-deployment subset is the items marked by code: timeouts, retries, idempotency, migrations, flags, and rollback. Those are also the items most likely to regress silently after launch, which is why they deserve a check on every pull request, not just on day one. See How to Assess Production Reliability Before Deployment for the per-change view.
The production readiness checklist for a backend service
Here is the full list, grouped by section. The downloadable versions add a one-line reason for each item and a suggested owner.
- A named owning team is recorded in the service catalog
- The service has a criticality tier that sets how strict this checklist is
- An on-call rotation of at least two people is live and has received a test page
- Escalation contacts for every critical dependency are documented
- On-call engineers know the incident process and severity levels
- SLIs are defined for availability and latency of user-facing operations
- SLOs and an error budget are agreed, with a documented response when the budget is spent
- Paging alerts fire on SLO burn rate or user-visible symptoms, not on CPU or single errors
- Every paging alert links to a runbook
- A dashboard shows traffic, errors, latency, and saturation per endpoint and per dependency
- Logs are structured and every line carries a request or trace id
- Logs contain no secrets, tokens, or unnecessary personal data
- Trace context (W3C traceparent) propagates across inbound calls, outbound calls, and messages
- Metrics carry the deployed version so a change can be correlated with a regression
- Every outbound call (HTTP, gRPC, database, cache, queue) has an explicit timeout
- Timeouts fit inside the caller’s deadline, leaving room for a retry if one is allowed
- Retries are capped, use exponential backoff with jitter, and only wrap idempotent operations
- Retries happen at one layer only
- Write endpoints that clients may retry are idempotent or accept an idempotency key
- On SIGTERM the service stops taking new work, drains in-flight requests, and exits within the grace period
- The readiness probe reflects ability to serve; the liveness probe checks only the process
- CPU and memory requests and limits are set from measured usage, and runtime memory fits inside the limit
- At least two replicas run across zones, with autoscaling bounds and a PodDisruptionBudget
- Every dependency is listed with its owner, SLO, and whether it is hard or soft
- Owners of shared dependencies have confirmed capacity for your expected peak
- Behavior is known and tested when each dependency is slow, down, or returning errors
- Soft dependencies have a fallback (cached value, default, or degraded response)
- Concurrency to each dependency is bounded by pools, bulkheads, or a circuit breaker
- Schema migrations are backward compatible with the previously deployed version
- Migrations have been run against production-sized data to measure lock time and duration
- Backups are automated and meet a stated recovery point objective (RPO)
- A restore has been tested end to end within the recovery time objective (RTO), and the date is recorded
- Data is classified, and sensitive data is encrypted at rest with audited access
- Secrets live in a secret manager, not in code, images, or committed config, and rotation has been tested
- Every endpoint requires authentication, and authorization is checked per resource
- The service runs with a least-privilege identity and restricted network access
- Dependency and container image scanning runs in CI with a policy for critical findings
- Inputs are validated, request sizes are limited, and traffic is encrypted in transit
- Deploys run from CI/CD with no manual steps on production hosts
- Rollout is progressive (canary or percentage) with automated checks on SLIs between stages
- Risky behavior ships behind a feature flag that can be turned off without a deploy
- Rollback has been exercised on this service and its duration is known
- Configuration changes get the same review and progressive rollout as code
- Expected peak traffic is estimated, including growth and planned events
- A load test at or above expected peak ran on a production-like environment
- The saturation point and the first resource to run out are known
- Inbound rate limits exist, and relevant cloud and third-party quotas are known
- A runbook exists for every paging alert: what it means, how to confirm, how to mitigate
- An architecture diagram and dependency list are current
- The API contract is documented and versioned
- Manual operations (replay, backfill, reprocess, data fix) are documented and safe to run twice
The sections below explain the items that most often get checked without being true, and what evidence to ask for.
| Section | Failure it prevents | Evidence to ask for |
|---|---|---|
| Ownership & on-call | Incidents with nobody to page or decide | Service catalog entry, on-call schedule, record of a test page |
| Observability | Users notice before you do | SLO doc, alert rules with runbook links, dashboard URL, a sample trace |
| Reliability | Hung threads, retry storms, duplicates, dropped requests on deploy | Client config showing timeouts, retry policy, probe definitions, shutdown test |
| Dependencies | A slow dependency takes the service down | Dependency list with tiers, fault injection or game day notes |
| Data | Rollback breaks on schema; data unrecoverable | Migration plan, last restore test date and duration |
| Security | Leaked secrets, broken authorization | Secret manager paths, authz tests, scan results |
| Deployment | A bad change reaches all users at once | Rollout config, flag names, last rollback drill |
| Capacity | Launch-day saturation | Load test report with the first saturated resource |
| Documentation | Slow, improvised incident response | Runbooks, architecture diagram, API contract |
Ownership and on-call: who gets paged?
Ownership is the first section because every other item depends on it. An alert without a recipient, a dependency without an escalation contact, or a runbook without a maintainer decays within months. Check three things concretely: the service catalog names a team (not a person who may leave), the on-call rotation includes at least two people so nobody is permanently on the hook, and a test page has actually reached the phone of whoever is on call this week.
Assign a criticality tier at this stage too. The tier decides how strict the rest of the checklist is: a tier-1 service might need multi-zone replicas, a 99.9% availability SLO, and a quarterly restore drill, while an internal reporting job might need only an owner, basic alerts, and a documented rerun procedure.
Observability: SLIs, SLOs, alerts, and correlation ids
The observability section fails most often in a specific way: the dashboards exist, but nobody has defined what “working” means, so alerts fire on CPU and miss the user-facing failure. Start from the user:
- SLIs: the proportion of valid requests that succeed, and the proportion served faster than a threshold, measured where users experience them (load balancer or API gateway, not inside a single pod).
- SLOs: a target for each SLI over a window, such as 99.9% of
POST /ordersrequests succeed over 30 days, and an agreed consequence when the error budget is spent. - Alerts: page on the rate at which the error budget is burning, which catches both sudden outages and slow degradation without paging on every blip.
The multiwindow burn-rate alert from the SRE workbook is a good default. With a 99.9% SLO, a burn rate of 14.4 sustained for one hour consumes 2% of a 30-day error budget; requiring the five-minute window too makes the alert reset quickly once the problem stops.
groups:
- name: orders-api-slo
rules:
- alert: OrdersApiErrorBudgetFastBurn
expr: |
(
sum(rate(http_requests_total{job="orders-api",code=~"5.."}[1h]))
/ sum(rate(http_requests_total{job="orders-api"}[1h]))
) > (14.4 * 0.001)
and
(
sum(rate(http_requests_total{job="orders-api",code=~"5.."}[5m]))
/ sum(rate(http_requests_total{job="orders-api"}[5m]))
) > (14.4 * 0.001)
for: 2m
labels:
severity: page
annotations:
summary: "orders-api is burning its 30-day error budget 14x too fast"
runbook_url: "https://runbooks.example.com/orders-api/error-budget-burn"
The runbook_url annotation is not decoration. It is the checklist item “every paging alert links to a runbook” enforced in the alert definition itself.
Logs and traces with correlation ids
During an incident the question is rarely “is something wrong?” It is “which request, which hop, which change?” That question is only answerable if every log line, span, and metric exemplar carries the same identifier. Use the W3C Trace Context traceparent header, which OpenTelemetry propagates by default, and make sure it crosses the boundaries teams forget: message queues, background jobs, and outbound calls made from thread pools.
{"ts":"2026-09-29T10:14:03.512Z", "level":"error",
"service":"orders-api", "version":"1.42.0",
"trace_id":"4bf92f3577b34da6a3ce929d0e0e4736", "span_id":"00f067aa0ba902b7",
"route":"POST /orders", "dependency":"payments-api", "duration_ms":2003,
"msg":"payment call timed out", "error":"context deadline exceeded"}
Note the version field. The first question in almost every incident is what changed, and a version label on logs and metrics answers it without opening the deploy history. The difference between observability and reliability matters here: observability shows you the failure once it happens, while most of the rest of this checklist is about making sure it does not.
Reliability: timeouts, retries, idempotency, shutdown, and probes
This is the section where production readiness is decided in code, and where it most often regresses after launch. Each item below has a concrete, reviewable form.
Timeouts on every outbound call
Many HTTP and database clients default to no timeout or a very long one. Without an explicit timeout, a slow dependency holds a thread, a connection, and memory for as long as it likes, and the service runs out of all three before the dependency ever returns an error. Set a connect timeout and a request timeout on every client, and make each timeout shorter than the deadline of the request that triggered it.
Retries with backoff, jitter, and a cap
Retries help with transient failures and hurt with everything else. The checklist item has four parts: retry only operations that are safe to repeat, cap the number of attempts, back off exponentially with random jitter so clients do not retry in lockstep, and retry at one layer only. The Amazon Builders’ Library covers the reasoning in depth.
import random, time
def call_with_retry(op, *, attempts=3, base=0.1, cap=2.0,
retryable=(TimeoutError, ConnectionError)):
for attempt in range(attempts):
try:
return op()
except retryable:
if attempt == attempts - 1:
raise
# full jitter: sleep a random time in [0, capped backoff]
time.sleep(random.uniform(0, min(cap, base * 2 ** attempt)))
The layering rule is the one reviewers miss. If the gateway, the service, and the client library each retry three times, one failing call becomes 27 attempts against a dependency that is already struggling. The sibling post How to Spot a Retry Storm Before Merge walks through the amplification math.
Idempotency for anything a client may retry
Timeouts plus retries guarantee duplicate requests. A client that timed out does not know whether the server processed the charge, so it tries again. Write endpoints that clients may retry need to be naturally idempotent (a PUT that sets state) or accept an idempotency key that the server stores alongside the result and checks before doing the work again. The same applies to queue consumers, which almost always see at-least-once delivery; see How to Prevent Duplicate Webhook Processing.
Graceful shutdown and health probes
Every deploy, scale-down, and node drain terminates pods. If the service does not shut down gracefully, each of those events drops in-flight requests, and the error rate spikes on every release in a way that looks like noise. In Kubernetes the sequence is: the pod is marked terminating, its preStop hook runs, then the container receives SIGTERM, and if it is still running when terminationGracePeriodSeconds (30 seconds by default) expires, it receives SIGKILL. Removing the pod from Service endpoints happens in parallel and can lag, which is why a short preStop delay is common. The pod lifecycle documentation describes the details.
ctx, stop := signal.NotifyContext(context.Background(),
syscall.SIGTERM, os.Interrupt)
defer stop()
srv := &http.Server{Addr: ":8080", Handler: mux,
ReadHeaderTimeout: 5 * time.Second}
go func() {
err := srv.ListenAndServe()
if err != nil && !errors.Is(err, http.ErrServerClosed) {
log.Fatal(err)
}
}()
<-ctx.Done() // SIGTERM received
ready.Store(false) // /readyz now returns 503
// 20 s drain fits inside 30 s grace minus the 5 s preStop sleep
shutdownCtx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
// stop accepting, wait for in-flight requests
if err := srv.Shutdown(shutdownCtx); err != nil {
log.Printf("forced shutdown: %v", err)
}
db.Close() // release the pool only after requests finish
Probes are the other half. The readiness probe answers “should this pod receive traffic right now?” and failing it removes the pod from load balancing. The liveness probe answers “is this process stuck beyond repair?” and failing it restarts the container. A liveness probe that checks the database is a classic mistake: when the database has a brief problem, every pod fails liveness at once and Kubernetes restarts the whole fleet, turning a dependency blip into a full outage. Use a startup probe for slow-starting services so liveness does not kill them during boot. The Kubernetes probe docs cover the fields.
spec:
replicas: 3
template:
spec:
terminationGracePeriodSeconds: 30
containers:
- name: orders-api
image: registry.example.com/orders-api:1.42.0
resources:
requests: { cpu: "500m", memory: "512Mi" }
limits: { memory: "512Mi" }
readinessProbe: # can it take traffic?
httpGet: { path: /readyz, port: 8080 }
periodSeconds: 5
failureThreshold: 2
livenessProbe: # process only, no dependencies
httpGet: { path: /livez, port: 8080 }
periodSeconds: 10
failureThreshold: 3
startupProbe: # shields slow boots
httpGet: { path: /livez, port: 8080 }
periodSeconds: 5
failureThreshold: 30
lifecycle:
preStop:
exec: { command: ["sleep", "5"] } # needs sleep in image
Resource limits and autoscaling
Set requests from measured usage under load, not from a template. For memory, the runtime has to fit inside the limit: a JVM, Node.js, or Python process that does not know its container limit will be OOMKilled rather than fail gracefully (see Kubernetes OOMKilled: Memory Leak or Memory Limit?). Many platforms set a memory limit but no CPU limit to avoid throttling; follow your platform’s policy and record why. Run at least two replicas spread across zones with a PodDisruptionBudget, and test that the autoscaler’s maximum is reachable within your quotas.
Dependencies: capacity, failure modes, and fallbacks
A service’s availability is bounded by its hard dependencies. List every dependency with its owner, its SLO, and whether the service can function without it. Then answer, for each one, what happens when it fails in each of the ways dependencies actually fail:
| Failure mode | What the service should do | How to verify |
|---|---|---|
| Down (fast errors) | Fail fast, return a clear error or fallback, do not retry beyond the budget | Block the dependency in staging; confirm error rate and latency of your own endpoints |
| Slow (no response) | Time out, shed load, stop sending traffic via a circuit breaker | Inject latency; confirm threads and pools do not saturate |
| Wrong (errors or bad data) | Validate responses, isolate the bad path, degrade the feature | Return malformed or partial responses from a stub |
| Overloaded by you | Bound concurrency and rate per dependency | Owner confirms capacity for your peak in writing |
Slow is the dangerous one. A dependency that is down returns errors in milliseconds, and the service stays healthy. A dependency that is slow holds every thread and connection that touches it, and the service fails for all traffic, including requests that never needed that dependency. Bulkheads (separate pools or concurrency limits per dependency) keep one slow neighbor from consuming everything. For soft dependencies, such as recommendations or analytics, decide on the fallback now: a cached value, a default, or omitting the feature.
Capacity is a two-way conversation. A new service can be the first real load test a shared database or internal API has ever seen. Ask the owners to confirm, with numbers, that your expected peak fits.
Data: backward-compatible migrations, backups, and tested restores
During a rolling deploy, old and new versions of the code run at the same time against one database schema. After a rollback, the old version runs against whatever schema the new version left behind. That is why the migration item says backward compatible with the previously deployed version, and why the safe pattern is expand and contract across several releases.
The subtle part is step 3. Rolling back from step 3 to step 2 code is safe because step 2 code still writes email and the backfill already copied existing rows. Also run each migration against production-sized data before release. Adding a column with a volatile default, creating an index without the database’s online option, or changing a column type can lock or rewrite a large table for far longer than staging suggests.
For backups, the item that matters is the restore. Record the recovery point objective (how much data you can lose) and recovery time objective (how long recovery may take), then restore a real backup into an isolated environment, time it, and write the date on the checklist. Backups that have never been restored fail in surprising ways: missing permissions, missing encryption keys, a restore that takes ten hours instead of one.
Security, deployment, and rollback: what must be true before launch?
The security items on a production readiness checklist are the ones that are cheap before launch and expensive after it. Secrets belong in a secret manager and are injected at runtime, never baked into images or committed config, and rotating one should be a tested procedure rather than an emergency. Every endpoint needs authentication, and authorization must be checked per resource: the most common serious API flaw is not a missing login, it is one user reading another user’s object by changing an id. Dependency and container scanning should run in CI with a clear rule for what blocks a release.
Deployment items are about limiting how many users a mistake reaches and how fast you can undo it:
- Progressive rollout: canary or percentage-based, with automated checks on the SLIs between stages, so a regression is caught at a small share of traffic.
- Feature flags for risky behavior, so it can be switched off in seconds without a deploy. Flags need owners and removal dates or they become permanent config debt.
- A tested rollback: actually roll back a release on this service before launch, and write down how long it took. The first rollback should not happen during an incident.
- Config is code: configuration changes get review and progressive rollout too, because a one-line config change can have the same blast radius as a code change.
Knowing which change to roll back is its own skill. Which Commit Caused the Production Incident? covers how to narrow it down, and Change Risk Management in Software Engineering covers matching rollout caution to how risky a change is.
Capacity testing and runbooks: the last mile before launch
A load test answers two questions: does the service handle expected peak on production-like infrastructure, and what breaks first when it does not? The second answer is more valuable. If the first resource to saturate is the database connection pool, you know which metric to alert on and what to scale. Test at or above your estimated peak, with realistic request mixes and data volumes, and include the dependencies or faithful stand-ins. Record inbound rate limits and the cloud and third-party quotas that could cap you first.
Runbooks close the loop. Each paging alert gets one, with three parts: what the alert means in plain language, how to confirm the impact, and the first mitigation steps with the exact commands or links. Add the manual operations the service will eventually need, such as replaying a dead-letter queue, backfilling a table, or reprocessing a day of events, and make sure they are safe to run twice. Keep an architecture diagram and dependency list next to the runbooks, because responders need to see the blast radius before they can contain it.
For each item, the reviewer should be able to click something: the alert rule, the dashboard, the load test report, the rollback drill notes, the pull request that added the timeout. If an item has no link, it is not done, it is believed done.
How to run a production readiness review with this checklist
A checklist becomes a review when someone outside the team looks at the evidence. A lightweight process that works for most teams:
- Set the tier. Agree on the service’s criticality and mark which items are required, recommended, or not applicable for that tier.
- Copy the checklist into the repository or launch ticket. The Markdown version works as a
docs/production-readiness.mdfile reviewed in a pull request; the CSV works in a tracker. - Attach evidence to every checked item. A link per item: dashboard, alert, test report, runbook, or pull request.
- Review with someone outside the team. An SRE, platform engineer, or senior engineer from another service asks “show me” for each item.
- Turn gaps into owned, dated work. Launch with known, accepted risks rather than unknown ones, and write the accepted risks down.
- Re-run it when things change. New tier, new owner, a major architecture change, or a serious incident are all reasons to review again.
The pre-deployment subset should not wait for a review meeting. Timeouts, retries, idempotency, migrations, and flags change in ordinary pull requests, which is where they should be checked. How to Review a Pull Request for Production Reliability Risks covers that side.
How Tomosu helps
The checklist above is useful without any tool; download it and use it. If you want a second opinion on the items that live in code, a Tomosu scan is an optional next step. Tomosu scans a repository and maps its code paths, then surfaces findings that line up with the reliability, dependency, and data sections of this checklist:
- Outbound calls without explicit timeouts, and retries that are unbounded, lack backoff, or stack across layers.
- Write paths that are retried but not idempotent, such as handlers and consumers that can process the same request twice.
- Remote calls inside transactions and other long holds that turn a slow dependency into pool exhaustion.
- Blast radius: which endpoints and jobs share a dependency, so a finding on a critical path is weighed differently from one in a batch job.
Findings roll up into the Production Reliability Index (PRI), which you can use as one input to the readiness review, alongside the evidence the checklist asks for. It does not replace the review: ownership, SLOs, restore tests, and load tests still need people to verify them.
Scan your repository with Tomosu →
Key takeaways
- A production readiness checklist verifies that a service can be operated, not just that it works. Scale the bar to the service’s tier.
- Every checked item needs evidence a reviewer can click. A checkbox without a link is an opinion.
- Alert on SLO burn rate and link every paging alert to a runbook. Put a trace id and version on every log line.
- Timeouts on every call, capped retries with backoff and jitter at one layer, and idempotency for anything retried.
- Liveness probes check the process, not its dependencies. Drain on
SIGTERMinside the grace period. - Migrations must work with the previous version of the code. A backup is only real after a timed restore.
- Roll out progressively, keep risky behavior behind flags, and rehearse the rollback before you need it.
Download it, adapt it, commit it. Both files carry the same 51 items; the CSV adds owner and status columns.
Frequently asked questions
What is a production readiness checklist?
A production readiness checklist is a list of operational requirements a service must meet, with evidence, before it serves customers. It covers ownership and on-call, observability, reliability patterns such as timeouts and retries, dependencies, data recovery, security, deployment and rollback, capacity, and runbooks. Teams use it as the basis of a production readiness review.
What should be on a production readiness checklist for a backend service?
At minimum: a named owner and live on-call rotation; SLIs, SLOs, and burn-rate alerts that link to runbooks; structured logs and traces joined by a trace id; explicit timeouts, capped retries with backoff and jitter, and idempotent write paths; graceful shutdown and correct health probes; known dependency failure modes and fallbacks; backward-compatible migrations and a tested restore; managed secrets and per-resource authorization; progressive rollout with a tested rollback; and a load test.
What is the difference between a service launch checklist and a pre-deployment checklist?
A service launch checklist is the one-time production readiness review before a service first takes real traffic, repeated when its tier, owner, or architecture changes. A pre-deployment checklist is the short set of checks every release passes, such as tests and scans, backward-compatible migrations, timeouts on new calls, feature flags for risky paths, and a known rollback. Most pre-deployment checks should be automated in CI/CD.
Who should own the production readiness review?
The service team owns the checklist and gathers the evidence. Someone outside the team, typically an SRE, platform engineer, or senior engineer from another service, reviews that evidence and asks to see it for each item. Gaps become tracked work with an owner and a date, and any risks accepted at launch are written down.
How do you keep a production readiness checklist from becoming a checkbox exercise?
Require evidence for every checked item: a link to the dashboard, alert rule, load test report, runbook, restore drill, or pull request. Scale required items to the service's criticality tier so teams are not asked to fake answers for items that do not apply. Check the items that live in code, such as timeouts, retries, and migrations, on every pull request rather than only at launch.
Should a Kubernetes liveness probe check the database?
Usually not. A liveness probe failure restarts the container, so if liveness depends on the database, a short database problem makes every pod fail at once and Kubernetes restarts the whole fleet. Liveness should check only that the process itself is responsive. Use the readiness probe to take a pod out of load balancing when it cannot serve.
How often should a production readiness checklist be re-run?
Re-run the full checklist when the service changes criticality tier, owner, or architecture, and after a serious incident. Some items need their own schedule: restore tests, rollback drills, and on-call test pages decay if they are only done once. The code-level items should be checked continuously as part of pull request review.
Is Mercari's production readiness checklist free to use?
Yes. Mercari publishes its production readiness checklist on GitHub under the MIT License. It is written for microservices at Mercari and Merpay, with emphasis on Go, Kubernetes, and Google Cloud, and scales required items by a production readiness level tied to each service's SLO. The checklist in this guide is independent and free to copy and adapt.
A service is ready for production when the team can show, item by item, that it can be owned, observed, recovered, and rolled back. Keep the checklist, and if you want a free look at the parts that live in code, run a scan. Assess your repository →