Skip to main content
Create your own
Lesson illustration

Ensuring Reliable Benchmarks Against Dead-Code Elimination and Input Reuse

Hello! Last lesson, you benchmarked a CLI processing function using a fresh input reader and output buffer for each operation. That gave you a sensible measurement boundary. This lesson checks two ways a benchmark can still answer the wrong question: the compiler may remove work whose result is unused, or repeated operations may receive input that an earlier operation has already consumed or changed.

By the end, you should be able to inspect a benchmark and explain why its intended work happens on every iteration. Allow about 35–40 minutes for the reading and examples.


Make sure there is work to measure

A Go benchmark is compiled with optimizations enabled. If a function’s result is discarded and the call has no observable effect, the compiler may remove the call. A very small ns/op figure can then describe the benchmark loop rather than the function you meant to study.

Read the Go team’s explanation of this problem and of the protection provided by testing.B.Loop.

More predictable benchmarking with testing.B.Loop

Read the Go blog’s explanation of how a benchmark can accidentally time an empty loop, and why Go 1.24 introduced a safer loop form.

In “Old benchmark loop problems,” start with the paragraph beginning “There is another, more subtle pitfall” and read through the explanation in “How testing.B.Loop helps.” Pay particular attention to the unused result in the example and to what the compiler is prevented from doing inside a b.Loop() loop.

For example, this traditional benchmark intends to measure a calculation:

func isOdd(x int) bool {
	return x&1 == 1
}

func BenchmarkIsOddWrong(b *testing.B) {
	for i := 0; i < b.N; i++ {
		isOdd(201) // Result discarded.
	}
}

Because nothing uses the answer, the compiler may eliminate the calculation. With Go 1.24 or later, prefer b.Loop(), and check a result after the loop:

func BenchmarkIsOdd(b *testing.B) {
	var last bool

	for b.Loop() {
		last = isOdd(201)
	}

	if !last {
		b.Fatal("isOdd(201) returned false")
	}
}

b.Loop() prevents the particular dead-code-elimination problem described in the Go blog by preventing calls from being inlined into its loop body. It also excludes work before and after the loop from timing. Checking last establishes that the benchmark produced the expected answer without putting that check in every timed operation.

This tiny function is an illustration, not a useful performance baseline: the overhead of benchmarking an extremely cheap operation can dominate its reported cost, and a constant input may not represent real use. The broader rule is to retain an observable result and use a loop structure that does not allow the work you care about to disappear. Neither measure proves that you chose representative inputs.

If you maintain a pre-1.24 b.N benchmark, do not assume an unused call is safe. Keep the result observable, use the appropriate timer controls for setup and checking, and treat unexpectedly tiny timings as a reason to investigate rather than as an immediate performance win.


Give each operation the intended input

The other failure mode does not require any compiler trickery. Some operations change their input; others consume a reader’s current position. Calling the function repeatedly may be perfectly real work, but not the same kind of work each time.

Eli Bendersky’s sorting example makes the distinction visible.

Common pitfalls in Go benchmarking - Eli Bendersky's website

Read Eli Bendersky’s “Benchmarking the wrong thing” example to see how an in-place operation changes the workload after its first iteration.

Under “Benchmarking the wrong thing,” begin with the paragraph introducing sorting in the slices package. Read through the comparison of the incorrect and corrected benchmarks. The examples use the older b.N loop; focus on the input-state problem, which also applies to b.Loop().

Suppose base contains unsorted integers. This benchmark is invalid if the goal is to measure sorting unsorted input:

data := slices.Clone(base)

for b.Loop() {
	slices.Sort(data)
}

Sort changes data in place. The first iteration sorts unsorted data; the remaining iterations sort an already sorted slice. b.Loop() ensures the calls are measured, but it cannot restore the input for you.

Here is a Go 1.24+ benchmark that measures sorting the same initially unsorted data on every operation:

package labels

import (
	"math/rand"
	"slices"
	"testing"
)

func BenchmarkSortUnsorted(b *testing.B) {
	rng := rand.New(rand.NewSource(42))
	base := make([]int, 4096)
	for i := range base {
		base[i] = rng.Intn(1_000_000)
	}

	var last []int

	for b.Loop() {
		b.StopTimer()
		data := slices.Clone(base)
		b.StartTimer()

		slices.Sort(data)
		last = data
	}

	if !slices.IsSorted(last) {
		b.Fatal("sort produced unsorted output")
	}
}

The fixed seed makes the fixture repeatable. Cloning gives every operation an unsorted slice; stopping the timer around that clone makes this specifically a measurement of sorting, not of sorting plus preparing its input. The final check confirms a completed result after timing ends. For very cheap operations, repeatedly stopping and starting the timer can itself make a benchmark less informative; choose a substantial workload or reconsider the measurement boundary.

There is no universal rule that copying must be excluded. If the real operation you want to compare includes making a copy, leave slices.Clone(base) inside the timed portion. State the question first, then place preparation on the appropriate side of the timer.


Apply both checks to the CLI benchmark

Return to BenchmarkProcessLabels from the previous lesson. Its input byte slice can be shared because ProcessLabels only reads those bytes. Its reader position cannot be shared: once a call has read through a bytes.Reader, the next call would begin at the end. Likewise, writing successive results into one bytes.Buffer without clearing it would accumulate output.

That is why the previous benchmark created both streams inside the loop:

for b.Loop() {
	in := bytes.NewReader(input)
	var out bytes.Buffer

	if err := ProcessLabels(in, &out); err != nil {
		b.Fatal(err)
	}
	last = &out
}

Creating fresh streams is part of that benchmark’s stated per-operation boundary. The final output check after the loop makes a completed result observable. Because b.Loop() has stopped timing at that point, the check does not become part of ns/op.

When reviewing any benchmark, ask two separate questions:

  1. Is the intended work actually executed? Look for discarded results and, with older b.N loops, opportunities for the compiler to eliminate calculations.
  2. Is each iteration doing the intended work? Identify anything the operation consumes, mutates, caches, or retains. Reset or reconstruct it when the workload requires a fresh state.

A benchmark can pass one check and fail the other. b.Loop() helps with compiler elimination, but it will happily time a thousand sorts of an already sorted slice. Fresh inputs fix that slice problem, but they do not make a discarded pure calculation observable in an older benchmark.


Takeaways

A trustworthy benchmark needs both observable work and the intended starting state for every operation. Use b.Loop() on Go 1.24+ to avoid the documented loop-elimination pitfall, check a completed result outside the timed loop, and deliberately refresh any input that an operation consumes or changes. Decide whether that refresh belongs inside the measurement based on what “one operation” is supposed to mean.

Next, you’ll read Go’s benchmark reports for time and allocations per operation. Those figures become useful only once you know what each reported operation actually did.

Can't find a good explanation? Sign up and we'll make it for you

Sign up