Hello! Welcome back to our exploration of modern quantum error correction.
In our last lesson, we built a crucial data generation pipeline using Stim. We learned how to simulate a noisy stabilizer code and extract pairs of (syndrome, pauli_error), which serve as the features and labels for a machine learning model.
Today, we will complete the second half of this process. We will take that generated data and use it to implement, train, and evaluate a decoder. This directly addresses our learning outcome: Implement, train, and evaluate a feed-forward neural network to predict the most likely Pauli error from a given error syndrome.
By the end of this lesson, you will have built your first end-to-end machine learning-based quantum error decoder. Given your background in machine learning engineering, this will involve familiar concepts—model building, training loops, and evaluation—but applied to the unique context of quantum error correction.
1. From Decoding to Classification
The core idea is to reframe the decoding problem as a supervised classification task. This is a common strategy in the literature, as it allows us to leverage powerful, well-understood ML tools.
The paper "Machine-Learning based Decoding of Surface Code..." provides a clear rationale for this approach and outlines the specific architecture we will be building.
Machine-Learning based Decoding of Surface Code ...
To understand the methodology, let's start by reading the section from the paper that maps the decoding problem to classification and details the neural network architecture.
Please read Section 4, 'Machine learning based syndrome decoding for surface code', up to and including subsection 4.1.4, 'Training our ML model'. Pay close attention to: The argument for why ML decoders can outperform traditional methods like MWPM (Section 4.1). The mapping of the problem to classification (end of Section 4.1.1). The specific architecture of the Feed-Forward Neural Network (FFNN), including the number of layers, nodes, and activation functions (Section 4.1.4).
As the paper describes, our task is a multi-label classification problem:
- Input Features (X): The syndrome, a binary vector where each element corresponds to a detector.
shape = (num_samples, num_detectors). - Output Labels (Y): The ground-truth Pauli error, a multi-hot encoded binary vector where each element corresponds to a possible physical error location and type.
shape = (num_samples, num_possible_errors). - The Model's Goal: Given a syndrome, predict a vector of probabilities, where each probability indicates the model's confidence that a specific Pauli error occurred.
2. Implementing the FFNN Decoder
We will now implement the FFNN architecture described in the paper you just read using TensorFlow/Keras. This framework is well-suited for this task and is also mentioned in the supplementary materials of another one of our resources ("A scalable and fast artificial neural network syndrome...", Appendix H).
Let's assume you have already generated and saved your data from the previous lesson. We'll load it and prepare it for training.
import numpy as np
import tensorflow as tf
from sklearn.model_selection import train_test_split
import stim
# --- Step 1: Load and Prepare Data ---
# Assume you have these files from the previous lesson's script
# Let's say we used a d=3, r=3 repetition code with p=0.01
# Load the generated data
syndromes = np.load('syndromes.npy')
pauli_errors = np.load('pauli_errors.npy')
# Split into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(
syndromes, pauli_errors, test_size=0.2, random_state=42
)
print(f"Training data shape: {X_train.shape}")
print(f"Training labels shape: {y_train.shape}")
# --- Step 2: Define the Model Architecture ---
# Based on the paper 'Machine-Learning based Decoding of Surface Code...'
num_detectors = X_train.shape[1]
num_possible_errors = y_train.shape[1]
# The paper uses 2 hidden layers with 32 and 16 nodes.
# This is a reasonable starting point for a small code.
model = tf.keras.Sequential([
tf.keras.layers.Input(shape=(num_detectors,)),
tf.keras.layers.Dense(32, activation='relu'),
tf.keras.layers.Dense(16, activation='relu'),
# Output layer with sigmoid for multi-label classification
tf.keras.layers.Dense(num_possible_errors, activation='sigmoid')
])
model.summary()
A quick note on the loss function: The paper mentions using Mean Squared Error (MSE). However, for a multi-label classification problem where each output is an independent probability (0 or 1), Binary Cross-Entropy (BCE) is the more conventional and theoretically sound choice. It treats each output neuron as a separate binary classifier. We will use BCE.
3. Training the Network
With the model defined, the next step is to compile and train it. We'll use the training parameters mentioned in the paper as a guide.
# --- Step 3: Compile and Train the Model ---
# The paper mentions SGD with a learning rate of 0.01.
# Adam is often a good default optimizer as well. Let's use it here.
optimizer = tf.keras.optimizers.Adam(learning_rate=0.01)
model.compile(
optimizer=optimizer,
loss='binary_crossentropy',
metrics=['binary_accuracy'] # Measures accuracy per output label
)
# The paper uses a large number of epochs, which is common.
# For demonstration, we'll use a smaller number.
# Batch size can be tuned; let's start with a reasonably large one.
history = model.fit(
X_train,
y_train,
epochs=50,
batch_size=1024,
validation_data=(X_test, y_test),
verbose=2 # Show one line per epoch
)
Test your understanding!
Imagine you switched from a distance-3 repetition code to a distance-5 surface code, which has more data qubits and stabilizers. How would the num_detectors and num_possible_errors change, and what effect would this have on the input and output layers of our neural network?
Show answer
Both num_detectors and num_possible_errors would increase significantly. A larger code has more stabilizers (and thus more detectors over the same number of rounds) and more physical qubits/gates where errors can occur.
This would directly change our network architecture:
- The input layer's size (
num_detectors) would become larger to accommodate the longer syndrome vector. - The output layer's size (
num_possible_errors) would also become larger to predict across a wider range of possible physical errors.
The hidden layers might also need to be made larger (more nodes) to handle the increased complexity of the mapping from syndrome to error.
4. Evaluating Decoder Performance
Standard ML metrics like binary_accuracy tell us how well the network predicts the exact physical error, but they don't tell us what we really care about in QEC: did we prevent a logical error?
A decoder can predict an incorrect physical error chain, yet still succeed if the combination of the true error and the predicted correction does not cause a logical flip. This is because different physical errors can be logically equivalent. The ultimate metric for a decoder is the Logical Error Rate (LER).
To calculate the LER, we perform the following steps for each sample in our test set:
- Get the syndrome and the true physical error
E_true. - Feed the syndrome to our trained model to get a predicted physical error
E_pred. - Calculate the residual error:
E_residual = E_true ⊕ E_pred(where ⊕ is bitwise XOR, as applying a Pauli twice cancels it). - Check if
E_residualconstitutes a logical error.
We can use Stim's Detector Error Model (DEM) to perform step 4. A DEM keeps track of which physical errors flip the logical observables.
# --- Step 4: Evaluate the Logical Error Rate (LER) ---
# We need the DEM from the circuit to check for logical errors.
# Let's assume you have a function to generate the circuit from the previous lesson.
# (This is a simplified version of the code from the last lesson)
def get_rep_code_circuit_and_dem(distance, rounds, noise_prob):
circuit = stim.Circuit.generated(
"repetition_code:memory",
distance=distance,
rounds=rounds,
after_clifford_depolarization=noise_prob
)
dem = circuit.detector_error_model(decompose_errors=True)
return circuit, dem
_, dem = get_rep_code_circuit_and_dem(distance=3, rounds=3, noise_prob=0.01)
# Get model predictions for the test set
# We convert probabilities to binary predictions with a 0.5 threshold
y_pred_probs = model.predict(X_test)
y_pred = (y_pred_probs > 0.5).astype(int)
# Now, calculate the LER
num_logical_errors = 0
num_test_samples = X_test.shape[0]
# Get the mapping from physical errors to logical observables
# This is a boolean matrix where rows are physical errors and columns are logical observables
error_to_logical_map = dem.get_logical_flip_table()
for i in range(num_test_samples):
# The actual physical error that occurred
true_error_vector = y_test[i]
# The physical error predicted by the decoder
predicted_error_vector = y_pred[i]
# The residual error after correction
residual_error_vector = (true_error_vector + predicted_error_vector) % 2
# Map the residual physical error to its effect on logical qubits
residual_logical_flips = (residual_error_vector @ error_to_logical_map) % 2
# Check if any logical qubit was flipped
if np.any(residual_logical_flips):
num_logical_errors += 1
logical_error_rate = num_logical_errors / num_test_samples
print(f"\nEvaluation on {num_test_samples} test samples:")
print(f"Number of logical errors: {num_logical_errors}")
print(f"Logical Error Rate (LER): {logical_error_rate:.4f}")
This LER is the single most important figure of merit for your decoder. In research papers, you will typically see plots of LER as a function of the physical error probability, which are used to find the code's threshold.
Conclusion
In this lesson, we completed the full workflow of building an ML-based quantum error decoder. We leveraged your expertise in ML to translate a physics problem into a familiar engineering task.
Key Takeaways:
- Decoding as Classification: The problem of inferring physical errors from a syndrome is a natural fit for multi-label classification.
- Implementation from Literature: We successfully implemented a simple Feed-Forward Neural Network based directly on an architecture proposed in a research paper.
- Training and Prediction: The training process follows standard ML practices, using data generated by a specialized quantum simulator (Stim).
- Evaluation with LER: The ultimate success of a decoder is not measured by prediction accuracy alone, but by its ability to prevent logical errors, quantified by the Logical Error Rate.
This simple FFNN is just the beginning. The 2D structure of codes like the surface code makes them amenable to more specialized architectures like Convolutional Neural Networks (CNNs), as explored in both of today's papers. These models can learn local correlations between syndrome bits, often leading to better performance.
Preview of the next lesson:
We have now covered the fundamentals of active error correction: we detect errors with stabilizers and then apply a corrective operation. In the next lesson, we will pivot to a radically different and fascinating paradigm: autonomous quantum error correction. We will explore how systems can be engineered to passively dissipate errors without any measurement or feedback, a key concept for building self-correcting quantum memories.