Hello! Last time, you learned to read ns/op, B/op, and allocs/op in a Go benchmark report—and to check what one benchmark operation actually includes. A single report, however, cannot tell you whether a small difference reflects a code change or ordinary measurement noise.
This lesson closes the performance-baseline module. You’ll collect repeated results for the same CLI benchmark, compare them with benchstat, and decide what the comparison does—and does not—establish. Allow about 35–40 minutes, including time to run the commands.
Collect comparable samples
benchstat compares benchmark output files, usually representing a before and an after version. Its summaries depend on having repeated measurements rather than one benchmark row per version.
benchstat command - golang.org/x/perf/cmd ...
Read the official benchstat documentation for its recommended sample count and an example of its comparison output.
In “Overview,” read the overview. Then, in “Example,” find the paragraph beginning “The table then compares the two input files.” Read the comparison explanation, focusing on the median, the comparison column, the sample count, and the ~ symbol.
Install the tool if you have not already:
go install golang.org/x/perf/cmd/benchstat@latest
Use the BenchmarkProcessLabels benchmark from the preceding lessons. Before changing the implementation, save ten measurements:
go test -run '^$' -bench '^BenchmarkProcessLabels$' -benchmem -benchtime=1s -count=10 > old.txt
Run your correctness tests, make one implementation change without changing the benchmark or its input, and run the correctness tests again. Then collect the after measurements from the same package, on the same machine:
go test -run '^$' -bench '^BenchmarkProcessLabels$' -benchmem -benchtime=1s -count=10 > new.txt
benchstat old.txt new.txt
Keep old.txt available when switching versions. If you have no change to compare yet, collect both files from the same version as a useful noise check.
The -count=10 flag gives benchstat ten results for each version. It is different from the large iteration count printed on each benchmark row: that count is the number of operations performed within one measurement. Keep the benchmark name, timed work, input size, flags, Go version, and machine conditions consistent. Otherwise, the files may describe different workloads rather than different implementations.
Interpret the comparison, not just the numbers
Suppose an illustrative sec/op row shows something like this:
old.txt new.txt vs base
ProcessLabels-8 1.200µ ± 3% 1.050µ ± 2% -12.50% (p=0.004 n=10)
The two central figures summarize the repeated time measurements; the accompanying percentages express uncertainty around those summaries. In the comparison column, a negative percentage for time per operation means the after version took less time. Here, n=10 refers to ten measurements from each file, not ten operations. A small p-value indicates that a difference this pronounced would be unusual if there were no systematic performance difference. It does not prove which code change caused it or guarantee the same result under every workload.
Now consider a smaller apparent improvement reported as:
ProcessLabels-8 1.200µ ± 4% 1.180µ ± 4% ~ (p=0.446 n=10)
The ~ means benchstat did not detect a statistically significant difference at its default threshold. It does not prove that the implementations have identical performance. The observed difference may be too small relative to the noise for these samples to resolve.
With -benchmem, inspect the allocation comparisons too. Faster time is not automatically an overall win if bytes or allocations per operation rise in a way that matters to your workload. Equally, a statistically detectable change can be too small to matter in practice. Judge direction, magnitude, uncertainty, and workload relevance together.
Keep noise from becoming a false conclusion
Repeated measurements help, but they do not repair an unfair experiment. Read the documentation’s measurement advice before treating a promising percentage as a result.
benchstat command - golang.org/x/perf/cmd ...
Return to the official documentation for practical guidance on measurement noise, run order, and repeated comparisons.
In “Tips,” read the measurement guidance. Pay particular attention to why the authors recommend interleaving before and after runs and choosing a sample count in advance.
For a first comparison, the two -count=10 commands above are straightforward. Their limitation is that all before runs happen first: background load or temperature could change before the after runs begin. For an important decision, precompile both versions’ benchmark binaries and alternate running them, appending each version’s results to its own file. This distributes changing machine conditions more evenly. Run on an otherwise idle machine where possible.
Choose your sample count before looking at the outcome—at least ten per version, with twenty preferable when feasible—and keep it fixed. If benchstat reports ~, do not repeatedly collect fresh batches until one happens to cross the significance threshold. At the default 0.05 threshold, chance alone can occasionally produce a “significant” result; searching for one makes that risk worse. Instead, report the inconclusive result, reduce identifiable noise, or plan a new comparison with more samples.
Finally, benchstat evaluates the measured benchmark, not the whole CLI. The timing and allocation boundary you established earlier still governs every claim you make from these files.
Takeaways
Use repeated, equivalent benchmark measurements as inputs to benchstat, then read its median summaries, percentage change, sample count, and significance indication together. A ~ is insufficient evidence of a difference—not proof of equality. Control the workload and machine conditions, and resist rerunning solely to obtain a favorable p-value.
You now have a defensible performance baseline. The next module turns to a different kind of trustworthiness: detecting shared-memory races before concurrent code is measured or optimized.