During an incident you usually have three disconnected facts: an error rate that went up, a pile of log lines, and a list of things that were deployed today. To correlate logs and traces with the code change that caused the error, you need join keys that connect them. A trace id joins a log line to a request. A version attribute joins the request to a build. A git SHA joins the build to a pull request.
To correlate logs and traces, put the trace id and span id from the W3C traceparent context into every structured log line, and tag all telemetry with service.name and service.version set to the deployed git SHA. Then an incident becomes a chain of joins:
- Metric → trace: an exemplar or a failing request gives you a
trace_id. - Trace → logs: filter every service’s logs by that
trace_id. - Trace → code: the error span names the service, operation, and stack frame.
- Code → version → change:
service.versiongives the SHA, the SHA range gives the pull requests. - Change → blast radius: the changed code’s callers show what else is affected.
Each join only works if something was set up before the incident. This guide covers the setup (propagation, log injection, version stamping), the step by step workflow during an incident, and the places where correlation silently breaks: async boundaries, message queues, and sampling.
How do you correlate logs and traces with a code change?
Correlation is a join, not a search. Searching all logs for “timeout” in the last hour finds thousands of lines from dozens of requests and several services, and you cannot tell which ones belong together. Joining on an id that every signal shares gives you exactly the lines that belong to one failing request, and then exactly the build and the change that served it.
The chain has five hops. Each hop has a key, and each key has to be emitted by your code or your platform before the incident starts.
The first two hops are the observability problem that OpenTelemetry’s logs specification calls log correlation: recording the execution context (trace id and span id) on log records so that logs and spans from the same request can be joined. The last three hops are the change management problem, the one postmortems record as the trigger and the root causes. Google’s example postmortem lists both fields separately, and for good reason: a traffic spike can be the trigger while a latent bug in a recent change is the root cause. Getting from one to the other quickly is what this workflow is for.
What does the traceparent header carry?
The W3C Trace Context recommendation defines two HTTP headers. traceparent carries the ids every system needs; tracestate carries optional vendor specific data. OpenTelemetry’s default propagator uses these headers, as its context propagation docs describe.
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
# ver trace-id (32 hex) parent-id (16 hex) flags
Three details matter for correlation:
- The trace-id never changes across the request. Every service that participates writes the same 32 hex character value, which is why it can join logs across services.
- The parent-id changes at every hop. It is the span id of the caller. Your service creates its own span with a new span id and sends that as the parent-id downstream.
- An all-zero trace-id or parent-id is invalid, and receivers must ignore the header. If your logs show
trace_id=00000000000000000000000000000000, the logging integration ran with no active span, not with a real trace.
How do you put the trace id in logs?
A trace id in logs is only useful if it is a separate, queryable field, not buried in the message text. Use structured (JSON) logs where you can, and name the fields consistently across services. The OpenTelemetry log data model uses TraceId and SpanId fields on each log record; the language integrations below expose them under their own names, so pick one convention and map everything to it in your log pipeline.
Java: the OpenTelemetry Java agent and MDC
The OpenTelemetry Java agent injects trace_id, span_id, and trace_flags into the MDC for Logback, Log4j 2, and Log4j 1. You only need to reference them in the layout:
<pattern>%d{ISO8601} %-5level [%thread] %logger{36}
trace_id=%X{trace_id} span_id=%X{span_id} trace_flags=%X{trace_flags} - %msg%n</pattern>
If you use a JSON encoder that writes MDC entries as fields, the ids appear as top level keys with no pattern change. Without the agent, copy the ids from the current span yourself at the edge of each unit of work, and clear them when it ends:
SpanContext ctx = Span.current().getSpanContext();
if (ctx.isValid()) { // skip the all-zero invalid context
MDC.put("trace_id", ctx.getTraceId());
MDC.put("span_id", ctx.getSpanId());
}
try {
chain.doFilter(request, response);
} finally {
MDC.remove("trace_id"); // MDC is thread local: always clear it
MDC.remove("span_id");
}
Python: the logging instrumentation
The OpenTelemetry logging instrumentation for Python adds otelTraceID, otelSpanID, otelServiceName, and otelTraceSampled to every LogRecord. Setting OTEL_PYTHON_LOG_CORRELATION=true also installs a default format that prints them:
import logging
from opentelemetry.instrumentation.logging import LoggingInstrumentor
LoggingInstrumentor().instrument(set_logging_format=True)
# or, for a custom format, reference the injected attributes yourself:
logging.basicConfig(format=(
"%(asctime)s %(levelname)s %(name)s "
"trace_id=%(otelTraceID)s span_id=%(otelSpanID)s "
"sampled=%(otelTraceSampled)s - %(message)s"
))
What a correlatable log line looks like
Whatever the language, aim for one shape. The ids join the line to a trace; the resource fields join it to a build:
{
"timestamp": "2026-09-29T14:17:03.412Z",
"level": "ERROR",
"message": "payment charge failed after 3 attempts",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"span_id": "53995c3f42cd8ad8",
"service.name": "checkout-api",
"service.version": "9f8e7d6",
"deployment.environment.name": "production",
"logger": "com.example.payment.PaymentClient",
"exception.type": "java.net.SocketTimeoutException"
}
Correlation fails most often on the log line that matters: an exception caught, logged with a custom message, and swallowed. If that code runs outside the request’s context (a callback, a scheduled retry), the line carries no trace id at all. Recording the exception on the active span as well (span.recordException(e) and an error status) means the trace shows the failure even when the log line is orphaned.
How do exemplars link metrics to traces?
Most incidents start with a metric: a latency percentile, an error ratio, a burn rate alert. Metrics are aggregates, so they have no trace id of their own. Exemplars close that gap. The OpenTelemetry metrics data model defines an exemplar as a sample recorded alongside a metric point: the measured value, the time, a set of filtered attributes, and optionally the trace_id and span_id of the request that produced it.
Exemplar support depends on the whole path: the SDK must record them (typically only for sampled spans), the exporter and the metrics backend must store them, and the dashboard must render them as links. Check each link before you depend on it. Without exemplars, the fallback is to query the tracing backend directly for error spans or slow spans in the same service and time window, which works but is slower and less precise.
How do you stamp the git SHA into every signal?
Hop 3 of the chain needs the running version on every span, metric, and log line. OpenTelemetry resource attributes are the right place because they are attached once per process and apply to all telemetry that process emits. The ones to set:
service.name: the logical service. Without it, telemetry from different services is hard to separate.service.version: the version string of the running build. Use the git SHA, or a release version that includes it (2.14.0+9f8e7d6). A semver string alone is not enough if you cannot map it back to a commit.deployment.environment.name:production,staging, and so on. Older SDKs and dashboards may still usedeployment.environment; check which one your backend expects.
OpenTelemetry also defines VCS attributes such as vcs.ref.head.revision and vcs.repository.url.full. At the time of writing they are not yet marked stable and are aimed mainly at CI/CD and repository telemetry, so treat service.version as the dependable join key and add VCS attributes if your backend uses them.
The SHA has to be captured at build time, because the running container has no .git directory. Pass it in once and expose it three ways: as a resource attribute, as an image label, and on an endpoint.
# build: docker build --build-arg GIT_SHA=$(git rev-parse --short=12 HEAD) .
ARG GIT_SHA=unknown
LABEL org.opencontainers.image.revision=$GIT_SHA
ENV GIT_SHA=$GIT_SHA \
OTEL_SERVICE_NAME=checkout-api \
OTEL_RESOURCE_ATTRIBUTES=service.version=$GIT_SHA
Set deployment.environment.name in the deployment manifest rather than the image, so the same image carries the same SHA through staging and production. When both the image and the manifest set OTEL_RESOURCE_ATTRIBUTES, the manifest value replaces the whole variable, so put service.version there as well, templated from the image tag. For compiled binaries the equivalent is a linker flag or build metadata. In Go, -ldflags "-X main.version=$(git rev-parse HEAD)" sets a variable, and modern Go toolchains also record vcs.revision in the build info that debug.ReadBuildInfo() returns when built from a git checkout.
{ "service": "checkout-api", "version": "9f8e7d6a41c2", "built_at": "2026-09-29T13:58:11Z" }
The endpoint answers the question every incident channel asks: “what is actually running right now?” The deploy log answers the other half: when each SHA started serving traffic, in which environment, and what it replaced. Together they give you the previous SHA..current SHA range that hop 4 needs.
Step by step: how to find the code change causing an error
With the keys in place, the workflow is mechanical. Run it in this order, and write down each key you find in the incident channel as you go.
- Start from the symptom and get one trace id. Open the alerting metric and pick an exemplar from the bad window, or query the tracing backend for error spans in the affected service. Take two or three trace ids, not one, so you do not chase an outlier.
- Open the trace and find the first failing span. Look for the deepest span with an error status or exception event. The top level span only tells you the request failed; the deepest error tells you where.
- Pull every service’s logs for that trace id. Filter by
trace_idacross all services, sorted by time. This shows retries, warnings, and swallowed errors that never became span events. - Read the resource attributes of the failing span. Note
service.name,service.version,deployment.environment.name, and the pod or host. Now you know which build served the failure. - Split the symptom by version. Group the error rate or latency by
service.version. If failures come from one version and not the previous one, the change is in that SHA range. If both versions fail equally, look outside the service: a dependency, data, config, or traffic. - List the changes in the range. Use the deploy log to get the previous and current SHA, then list the commits and merged pull requests between them, filtered to the code location from step 2.
- Check the blast radius before you act. Find the callers of the changed code and what else shares it. That tells you whether rollback is safe, what else might be broken, and whether other services need attention.
- Record trigger and root cause separately. In the postmortem, the trigger might be traffic or a dependency slowdown; the root cause is often the change that made the system fragile to it.
Step 6 in practice, once you have both SHAs and the code location from the failing span:
# Everything that shipped in the bad deploy
git log --oneline a1b2c3d..9f8e7d6
# Only merge commits (one per pull request, for merge-commit workflows)
git log --merges --oneline a1b2c3d..9f8e7d6
# Only commits that touched the failing code path from the error span
git log --oneline a1b2c3d..9f8e7d6 -- src/main/java/com/example/payment/
# History of the specific function named in the stack frame
git log a1b2c3d..9f8e7d6 \
-L :charge:src/main/java/com/example/payment/PaymentClient.java
If the range is large and the code location does not narrow it enough, bisect the range by deploying or canarying intermediate SHAs, or by reverting candidate pull requests one at a time. The sibling post Which Commit Caused the Production Incident? covers that narrowing in depth. For step 7, How to Assess the Blast Radius of a Code Change describes how to trace callers, shared resources, and downstream consumers.
“Errors started five minutes after the 14:12 deploy” is the most common and the weakest argument in an incident channel. Several things usually change in the same window: a config push, a feature flag, a dependency release, a traffic shift. The version split in step 5 and a failing span inside changed code are what connect the error to a specific change.
What breaks correlation between logs, traces, and code?
Correlation fails quietly. Nothing errors when a trace id goes missing; you just find, mid-incident, that the log line you need has an empty field or a trace id that points nowhere. These are the usual causes.
Async boundaries inside a service
Trace context lives in a context object bound to the current thread or coroutine. Hand work to an executor, a scheduled task, a reactive pipeline, or a callback, and the new execution may start with no active span. The logs from that work then have empty trace fields, or worse, a new trace id unrelated to the request. Auto-instrumentation covers many common executors and frameworks, but custom thread pools, hand-rolled queues, and some async libraries need explicit wrapping, for example with Context.current().wrap(runnable) in the OpenTelemetry Java API. The MDC has the same problem: it is thread local, so values do not follow the work to another thread unless something copies them.
Message queues and batch consumers
HTTP instrumentation propagates traceparent automatically. Message brokers are hit and miss. The producer has to write the context into message headers or attributes, and the consumer has to extract it. If either side is uninstrumented, the consumer starts a fresh trace and the join from the producer’s request to the consumer’s failure is gone. Batch consumers add a second problem: one poll can contain messages from many different traces, so there is no single parent. OpenTelemetry’s answer is span links, which let one consumer span reference several producer contexts. Your log lines, though, can only carry one trace id at a time, so log per message inside the batch loop, with that message’s context active.
Sampling
Head sampling decides at the start of a trace whether to record it, and the decision travels in the trace flags. Logs are usually not sampled. So with a 10% head sampler, most log lines carry a valid trace id for a trace that does not exist in your backend. That is expected, not a bug, and it is why the trace_flags or otelTraceSampled field is worth logging: it tells the responder in advance whether a trace will be there. Tail sampling, which decides after the trace completes, can keep all error traces, which is usually what incidents need.
Other quiet failures
- A proxy, gateway, or legacy service that drops unknown headers, which splits one request into two traces.
- Multiple propagation formats (for example W3C and a vendor or B3 format) configured differently across services, so one side injects a header the other does not read.
- Log fields named differently per service (
trace_id,traceId,otelTraceID) with no mapping in the log pipeline, so one query cannot join them. service.versionset tolatestor to a mutable tag, which makes hop 3 impossible.- Deploy logs that record the image tag but not the SHA, or record the time the pipeline finished rather than the time traffic shifted.
Evidence table: what each signal proves
Each piece of evidence answers one question and leaves others open. Be explicit in the incident channel about which question a given piece of evidence answers.
| Evidence | Join key | What it proves | What it does not prove |
|---|---|---|---|
| Metric spike | time, service.name | Something is wrong, and when it started | Which requests, which code, or which change |
| Exemplar | trace_id | A specific request contributed to the spike | That the request is typical; take several |
| Error span | trace_id, span_id | Where in the call path the failure happened | Why it failed, or since when |
| Correlated logs | trace_id | The sequence of events, retries, and swallowed errors in that request | Anything about requests whose trace ids you did not pull |
| Resource attributes | service.version | Which build served the failing request | That the build is the cause |
| Version split | service.version | Failures are specific to one version | Which commit in the range is responsible |
| Commit range + code location | SHA range, file, function | Which change touched the failing code | That the change is wrong rather than exposed by other factors |
| Blast radius | call graph | What else depends on the changed code | That those dependents are failing now |
The last row is where most incident tooling stops and most follow-up work starts. Knowing that the change in PaymentClient.charge caused the checkout errors does not tell you whether the refund worker, which calls the same client, is quietly failing too. That is a question about code, not telemetry.
Telemetry tells you where the request failed. Only the code and its history tell you what changed there and what else depends on it.
How Tomosu helps
Tomosu works from the code side of this chain. It scans repositories and pull requests, maps code paths and their callers, and weighs findings by blast radius. That is hops 4 and 5: from a change to the code it touches, and from that code to everything that depends on it. Runtime Signals is one of the indexes in the Production Reliability Index, alongside Fragility, Drift, Code Volatility, and Deployment Velocity, so the score reflects how code behaves in production and not only how it reads.
The direction we are building toward is the full bridge described in this post:
- Runtime evidence to code location: a failing span, exception, or log line resolved to the service and the function it came from.
- Code location to change: the deployed version mapped to its commit range, and the pull requests in that range that touched the failing path.
- Change to blast radius: the callers, shared resources, and downstream services of the changed code, so the responder knows what else to check before choosing between rollback and fix.
- Back into review: the same map applied before merge, so a change that touches a high blast radius path is flagged when it is a pull request, not when it is an incident. See Pre Merge Reliability Analysis.
Today, the concrete starting point is the code map and the PRI: scan a repository to see which paths carry the most risk and what depends on them, so the last hops of the join chain are ready before the next incident needs them.
Scan your repository with Tomosu →
Key takeaways
- To correlate logs and traces, put
trace_idandspan_idinto every structured log line as separate fields with one naming convention across services. - The W3C
traceparentheader carries version, a 32 hex trace-id, a 16 hex parent-id, and flags. The trace-id is the join key; the flags say whether a trace was recorded. - Set
service.versionto the git SHA (or a version that maps to it) as a resource attribute, and expose it on an image label and a/versionendpoint. - Use exemplars to go from a metric spike to a specific trace, then from the trace to its logs and its version.
- Split the symptom by
service.versionbefore blaming a deploy. A time correlation is not evidence. - Async boundaries, message queues, and sampling break correlation silently. Test propagation through each of them before an incident does.
- The chain does not end at the commit. Check the blast radius of the change before you decide how to fix it.
Frequently asked questions
How do I correlate logs and traces?
Record the active trace id and span id on every log line as structured fields, using OpenTelemetry instrumentation or your logging library’s context (for example the MDC in Java). Propagate trace context between services with the W3C traceparent header. Then filter logs by the trace id from a failing trace to see every log line from that request across all services.
How do I add the trace id to logs?
With the OpenTelemetry Java agent, trace_id, span_id, and trace_flags are injected into the MDC for Logback and Log4j, so you reference them in the log pattern or JSON encoder. In Python, the OpenTelemetry logging instrumentation adds otelTraceID and otelSpanID to each log record. Without instrumentation, read the current span context and copy its ids into your logging context at the start of each unit of work.
What is the format of the traceparent header?
traceparent has four dash separated lowercase hex fields: a 2 character version (currently 00), a 32 character trace-id, a 16 character parent-id that is the caller’s span id, and 2 characters of trace flags, where 01 means sampled. An all-zero trace-id or parent-id is invalid. Example: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01.
Why do my logs have a trace id but the trace is missing?
Usually sampling. Head sampling decides at the start of a request whether to record the trace, but log lines are written with the trace id either way, so most ids point at traces that were never exported. Log the sampled flag so responders know in advance, and consider tail sampling that keeps all error traces.
How do I find the code change causing an error in production?
Get a trace id for a failing request, find the deepest error span, and read its service.version resource attribute. Split the error rate by service.version to confirm the failure is specific to one build. Then list the commits and pull requests between the previous and current deployed SHAs, filtered to the file or function named in the error span.
What are exemplars in metrics?
An exemplar is a sample measurement stored with a metric data point, together with the trace id and span id of the request that produced it. Dashboards can render exemplars as links, so you can go from a latency or error spike directly to a representative trace instead of searching for one.
Why does the trace break at a message queue or background job?
Trace context is not carried automatically across every boundary. For a queue, the producer must write traceparent into message headers and the consumer must extract it; batch consumers need span links because one batch contains many traces. For thread pools and background jobs, the context must be passed to the new thread, which auto-instrumentation does for common executors but not always for custom ones.
What should service.version contain?
A value that maps unambiguously to a commit: the git SHA, or a release version that includes it such as 2.14.0+9f8e7d6. Avoid mutable values like latest. Set it at build time and pass it as an OpenTelemetry resource attribute so every span, metric, and log line from the process carries it.
An incident is solved when the trace id, the version, and the change line up. Tomosu maps the code side of that chain, from a change to everything it touches, before the next incident asks for it. Assess your repository →