Hello! Last time, you wrote tests that protect CleanLabels against behavior changes. Now you can measure it without mistaking benchmark preparation for the cost of the function itself.
In this lesson, you’ll write a benchmark whose timed work is one call to CleanLabels per operation. You’ll prepare its input before timing starts, check the result after timing stops, and see how to express the same boundary on Go versions before 1.24. Allow about 35–40 minutes, including the short readings and a run of the code.
Decide what one operation means
A benchmark answers a question about a defined unit of work. Here, one operation is calling CleanLabels on a prepared slice. Constructing that slice is benchmark setup; it is not part of the operation we’ve chosen to measure. The slice and map that CleanLabels creates inside the call, however, are part of that operation and must remain timed.
Go discovers functions named Benchmark... in _test.go files. Unlike a test, a benchmark accepts *testing.B; you run it with go test -bench. In Go 1.24 and later, b.Loop() provides a particularly clear timing boundary: work before the loop is excluded, the loop body is timed, and work after the loop is excluded.
Read the Go testing package documentation for the benchmark function shape and the timer behavior we’ll use.
In the “Benchmarks” subsection, begin at benchmark basics, then inspect the immediately following BenchmarkBigLen example; stop after that example. In the B.Loop API entry, read timer boundaries. Focus on when the timer starts and stops, rather than on the example’s particular operation.
The distinction is more precise than “put the important code somewhere in a benchmark.” Code inside the loop determines the reported time per operation. Moving a required part of that operation outside the loop would make the result look faster without making the operation faster.
Benchmark the code you already tested
Keep labels.go and its correctness tests from the previous lesson. Add the following to a new file, labels_bench_test.go, in the same package. This version requires Go 1.24 or later; a version for older Go appears below.
package labels
import (
"slices"
"testing"
)
func BenchmarkCleanLabels(b *testing.B) {
// Setup: create the caller's input once, before timing starts.
input := make([]string, 1024)
for i := range input {
if i%4 == 0 {
input[i] = " RUST "
} else {
input[i] = " Go "
}
}
want := []string{"rust", "go"}
var got []string
// Timed work: one CleanLabels call per benchmark operation.
for b.Loop() {
got = CleanLabels(input)
}
// Post-benchmark check: the timer has stopped.
if !slices.Equal(got, want) {
b.Fatalf("CleanLabels() = %q, want %q", got, want)
}
}
The first input is " RUST ", so the expected order is rust, then go. The fixture-building loop runs before the first call to b.Loop() and is not included in the reported time. The result check runs after b.Loop() finishes and is not included either. If setup had acquired a resource that needed closing, you could likewise close it after this loop without timing the close.
Notice what stays inside the timed call: normalization, deduplication, and the allocations made by CleanLabels. Moving any of those into setup would change the question the benchmark answers. The post-loop check is a useful sanity check on the last result, but it does not replace the table-driven tests that check the function’s wider contract.
Reusing this input is appropriate because CleanLabels is specified not to modify it. A function that consumes or mutates its input may face a different task on its second iteration. We’ll examine that benchmark pitfall in a later lesson; for now, the key is to make every timed call represent the same intended operation.
Run it and read the boundary
First, make sure the behavior guard still passes:
go test ./...
Then run just this benchmark:
go test -run '^$' -bench '^BenchmarkCleanLabels$' -benchtime=1s .
The -run '^$' flag skips ordinary tests during the benchmark run; you have already run them separately. -bench selects the named benchmark. In its output, the iteration count tells you how many calls were timed, and ns/op reports elapsed time per loop-body operation—here, per CleanLabels(input) call. It is not the time to build the 1,024-element fixture or perform the final check.
If you change the fixture size, you change the work CleanLabels receives on every iteration, even though constructing the fixture remains untimed. Keep the input shape fixed when you later compare two implementations.
If your Go version predates 1.24
The previous lesson’s code works on Go 1.22 or newer, but b.Loop() was added in Go 1.24. Check your version with go version. On an older version, keep the same setup, expected result, and final check, but replace the b.Loop() loop with this block:
b.ResetTimer()
for i := 0; i < b.N; i++ {
got = CleanLabels(input)
}
b.StopTimer()
Use either this b.N loop or b.Loop() in a benchmark, not both. b.ResetTimer() discards time spent in setup; b.StopTimer() excludes the result check and any subsequent cleanup. Without those calls, setup or teardown in a b.N-style benchmark can contaminate the measurement.
The Go blog contrasts the manual timer calls with the newer loop and describes a case where preparation must happen on every iteration. Read it with the CleanLabels fixture in mind: its input can be prepared once only because the function leaves that input unchanged.
More predictable benchmarking with testing.B.Loop
Read the Go blog’s comparison of manual timing and b.Loop(), then its example of input that must be prepared repeatedly.
In “Old benchmark loop problems,” start at manual timing. Then, later in the article, find the BenchmarkSortInts example immediately before the paragraph beginning “In this example”; read the example and its explanation. Focus on why sorting cannot repeatedly use an already-sorted input, and why timer control is still needed when untimed preparation occurs inside the loop.
Takeaways
A trustworthy benchmark names its operation and places the timer around exactly that work. For CleanLabels, fixture construction belongs before the loop, the function call belongs inside it, and result checking belongs after it. Go 1.24’s b.Loop() handles those outer timing boundaries automatically; older b.N benchmarks need b.ResetTimer() and b.StopTimer().
Next, you’ll use the same principle to benchmark an HTTP handler in memory, separating handler work from the machinery used to supply a request.