Skip to main content
Create your own
Lesson illustration

Benchmarking a CLI’s Core with In-Memory I/O

Hello! Last lesson, you called an HTTP handler directly with an in-memory request and response recorder. A command-line program can be measured at a similar boundary: supply its input and capture its output without launching a process.

This lesson uses the CleanLabels function from earlier lessons to build a small CLI processing core. You’ll benchmark the work between reading newline-separated labels and writing JSON, while leaving terminal, file, and process-startup costs out of the result. Allow about 35–40 minutes to read and try the example.


Put a boundary around the CLI’s work

A CLI often obtains input from os.Stdin and sends output to os.Stdout. Those are useful choices when running the program, but they are poor fixtures for a repeatable benchmark. Instead, make the processing function accept an io.Reader and an io.Writer. The CLI entry point can pass it standard input and output; a benchmark can pass it memory-backed implementations.

The bytes package provides both sides of that arrangement. Read the relevant parts of its official documentation before looking at the benchmark.

bytes package - bytes - Go Packages

The Go bytes package documentation explains the in-memory reader and output buffer used below. Focus on which object is consumed by reading and which object grows when written to.

Under “Types,” read the descriptions of type Reader and func NewReader, beginning with the reader overview and continuing through its constructor. Then read type Buffer’s type description. In func (*Buffer) Write, start at “Write appends the contents” and read through the write result; in func (*Buffer) String, read from “String returns the contents” through the returned value.

The important asymmetry is that reading advances a reader’s position, while writing adds to a buffer. A benchmark must therefore give every operation an input reader at the beginning of the data and an empty place for its output.


Write the processing core

Keep the existing CleanLabels([]string) implementation. Add labels_cli.go to the same labels package:

package labels

import (
	"bufio"
	"encoding/json"
	"io"
)

// ProcessLabels reads one label per line and writes the cleaned labels as JSON.
func ProcessLabels(in io.Reader, out io.Writer) error {
	scanner := bufio.NewScanner(in)

	var labels []string
	for scanner.Scan() {
		labels = append(labels, scanner.Text())
	}
	if err := scanner.Err(); err != nil {
		return err
	}

	return json.NewEncoder(out).Encode(CleanLabels(labels))
}

For input containing the lines " RUST ", " Go ", and " RUST ", the expected output is ["rust","go"] followed by a newline. The function reads and collects the lines, calls CleanLabels, encodes the result, and writes it. Those steps are the processing core for this example; none should be performed in benchmark setup.

There is no dependency on os.Stdin or os.Stdout inside ProcessLabels. A real CLI entry point can call it with those streams and handle the returned error. That separation lets the benchmark measure the same processing function the CLI uses, without measuring process startup or physical I/O.


Benchmark one complete input

Add labels_cli_bench_test.go. As in the previous lesson, this b.Loop() version requires Go 1.24 or later.

package labels

import (
	"bytes"
	"testing"
)

func BenchmarkProcessLabels(b *testing.B) {
	// Construct the workload once; ProcessLabels does not modify these bytes.
	input := []byte(" RUST \n  Go  \n RUST \n")
	const want = "[\"rust\",\"go\"]\n"

	var last *bytes.Buffer

	for b.Loop() {
		in := bytes.NewReader(input)
		var out bytes.Buffer

		if err := ProcessLabels(in, &out); err != nil {
			b.Fatal(err)
		}
		last = &out
	}

	// b.Loop has stopped timing. Check a completed result.
	if last == nil || last.String() != want {
		b.Fatalf("output = %q, want %q", last, want)
	}
}

One operation here means processing the entire three-line input and capturing its JSON output. Each iteration creates a new reader because the preceding call consumed its reader. It also creates a new buffer so output cannot accumulate across iterations. The input bytes are shared safely: this function reads them but does not change them.

Preparing the byte slice and checking the final answer are outside the timed loop. Creating the reader and buffer is inside it because supplying fresh streams is part of this benchmark’s per-operation boundary. The final check catches an obviously invalid result; it does not replace tests for other inputs or error cases.

Run the existing tests, then the selected benchmark:

go test ./...
go test -run '^$' -bench '^BenchmarkProcessLabels$' -benchtime=1s .

On Go versions before 1.24, use the earlier b.N pattern in place of b.Loop(): call b.ResetTimer() after preparing input, run the body b.N times, then call b.StopTimer() before checking last.


Know what the number answers

The reported ns/op estimates the time for one in-memory, three-line processing operation. It includes scanning, label collection, CleanLabels, JSON encoding, writing to the buffer, and creation of the per-operation reader and buffer. It excludes launching a CLI process and reading from or writing to an actual terminal or file.

That distinction matters when choosing a benchmark. The earlier CleanLabels benchmark isolates the transformation itself; this one measures how the CLI’s processing path uses that transformation. Neither number alone predicts end-to-end runtime for a large file. Keep the input representative of the workload you want to improve, and keep its contents and the operation boundary unchanged when comparing future versions.


Takeaways

Accepting io.Reader and io.Writer lets a CLI’s real processing code run entirely in memory. Give each benchmark iteration a reader at the start of the input and a fresh output buffer, then verify a completed output after timing stops. The resulting baseline measures the processing core, not the surrounding operating-system I/O.

Next, you’ll examine how compiler elimination and unintended input reuse can make even a well-structured benchmark report a misleading result.

Can't find a good explanation? Sign up and we'll make it for you

Sign up