Skip to main content
Create your own
Lesson illustration

Quantization and Bit Depth: Foundations of Audio Fidelity

Hello! Welcome to the third lesson in our course on audio AI.

In the previous lesson, we explored the Nyquist-Shannon Sampling Theorem, which explains how we can discretize the continuous time axis of an analog signal into a series of discrete samples. You learned that if we sample at a rate greater than twice the highest frequency in the signal, we can perfectly reconstruct the original wave. This answers the question of how often to measure the wave's amplitude.

Now, we must address the second part of the analog-to-digital conversion puzzle: what about the amplitude values themselves? An analog signal has a continuous, infinite range of possible amplitude values at any given point. To store these on a computer, we must also discretize the amplitude axis.

This lesson covers the process of quantization and the crucial role of bit depth in determining the fidelity of digital audio. We will explore how mapping continuous amplitudes to a finite set of values introduces a specific type of error and how this directly impacts the dynamic range and perceived quality of the sound.


1. From Continuous Amplitude to Discrete Levels

In the last lesson, we took samples at discrete points in time. However, the value of each sample's amplitude is still a real number with potentially infinite precision (e.g., 0.5, 0.51, 0.512, 0.5123...). A computer cannot store numbers with infinite precision. We need to map this continuous range of amplitudes to a finite set of discrete levels. This process is called quantization.

Think of it as rounding. If you can only use integers, a value like 3.7 gets rounded to 4, and 3.2 gets rounded to 3. Quantization applies this concept to the amplitude of an audio signal.

Let's start with a video that introduces this concept and clarifies its role in the digitization process.

5. Quantization - Digital Audio Fundamentals

This video, 'Quantization' by Akash Murthy, explains why quantization is necessary for the amplitude axis, just as sampling is for the time axis. It defines quantization as the process of mapping analog signal values to a limited set of discrete values.

Watch the following two clips: 00:40 - 02:28: This section introduces the need to limit the resolution of the amplitude measurement and the trade-offs involved. 04:34 - 05:40: This part provides a clear visual demonstration of the quantization process, showing how sample values are 'snapped' to the nearest discrete level.

The core idea is that each sample's continuous amplitude is approximated by the nearest available discrete level. The number of available levels is determined by the bit depth.

2. Bit Depth: The Language of Levels

Your computer science background gives you a head start here. Digital information is stored in bits. The number of bits used to represent the amplitude of a single sample is the bit depth.

The relationship is straightforward: an n-bit number can represent distinct values.

  • 8-bit audio uses levels.
  • 16-bit audio (CD quality) uses levels.
  • 24-bit audio (studio quality) uses levels.

Most audio signals are signed, meaning they represent both positive and negative amplitudes (the wave above and below the zero line). For this, computers use a representation called two's complement, which allows an n-bit integer to represent values from to .

The following video explains how bit depth defines the number of quantization levels.

6. Bit Depth - Digital Audio Fundamentals

The follow-up video, 'Bit Depth', connects the concept of quantization levels directly to binary representation. It shows the exponential growth in representable values as bit depth increases.

Watch from the beginning to 01:53. This segment clearly explains how the number of available states grows exponentially with each added bit.

This image provides a clear visual comparison. On the left, a low bit depth results in a coarse, "blocky" approximation of the analog wave. On the right, a higher bit depth allows for a much finer and more accurate representation because there are more discrete amplitude steps available.

Quantization and Bit Depth Comparison
This image illustrates the effect of bit depth on quantization. The left panel shows a low bit depth with few quantization levels, resulting in a coarse approximation of the original analog wave (green). The right panel shows a high bit depth with many levels, leading to a much more accurate digital representation.

3. The Inevitable Consequence: Quantization Error

Quantization is inherently a lossy process. The difference between the original analog amplitude and the rounded, quantized amplitude is called quantization error.

This error isn't random. It manifests as a distinct type of distortion or noise that is added to the signal. We call this quantization noise. The lower the bit depth, the larger the rounding errors, and the more audible the noise becomes.

This is best understood by listening. The next resource provides a powerful demonstration.

6. Bit Depth - Digital Audio Fundamentals

This is the most crucial part of the lesson. The same 'Bit Depth' video will now be used to demonstrate the audible effects of severe quantization. It takes a clean audio track and reduces its bit depth, first to 4-bit and then to 8-bit, and even isolates the resulting quantization noise.

Watch from 03:56 to 11:50. Pay close attention to: The 'minecraft-like' blocky waveform at 4-bit (06:49). The horrendous noise introduced in the 4-bit audio example (07:03). How the isolated noise signal is extracted (07:49). How much less noticeable the noise is at 8-bit, but still present (10:50).

As you heard, with very low bit depth (4-bit), the noise is overwhelming. With a higher bit depth (8-bit), it becomes much less prominent but is still correlated with the signal, especially noticeable in quieter parts. This raises the central question: how do we measure this effect on audio fidelity?

4. Quantifying Fidelity: Dynamic Range and SQNR

The primary measure of fidelity affected by bit depth is dynamic range.

Dynamic Range is the ratio, measured in decibels (dB), between the loudest possible undistorted signal and the quietest perceptible signal. In digital audio, the "quietest" signal is limited by the noise floor, which is created by quantization error.

A higher bit depth results in smaller quantization steps, which lowers the noise floor and therefore increases the dynamic range.

The Math Behind the Noise

We can formalize this relationship using the Signal-to-Quantization Noise Ratio (SQNR). For a sine wave quantized with n bits, the theoretical maximum SQNR is given by:

Signal to Quantization Noise Ratio Derivation
This image shows the mathematical derivation for the Signal-to-Quantization Noise Ratio (SQNR), which quantifies the quality of quantization. It demonstrates that the SQNR in decibels is linearly proportional to the bit depth (n).

This formula gives rise to a very famous rule of thumb in digital audio:

Each additional bit of depth adds approximately 6 dB to the dynamic range.

Let's walk through a simplified derivation for this rule, as presented in the "Digital Signals Theory" text.

Quantization — Digital Signals Theory

This textbook excerpt by Brian McFee provides the formal definitions and derivations that connect bit depth to dynamic range.

Please read section '2.4.4. Dynamic range'. It cleanly derives the '6 dB per bit' rule and provides a helpful table showing the dynamic range for common bit depths.

Let's summarize the derivation:

  1. The dynamic range is the ratio of the maximum possible value () to the minimum non-zero value ().
  2. For an n-bit signed integer, the maximum value is proportional to , while the smallest non-zero value is proportional to 1.
  3. The ratio is therefore .
  4. Converting this ratio to decibels for an amplitude-like quantity gives:

This simple approximation tells us the theoretical dynamic range for common bit depths:

  • 8-bit: dB (Very limited, noisy)
  • 16-bit (CD): dB (Good for listening, as it exceeds the dynamic range of most environments)
  • 24-bit (Studio): dB (Exceeds the dynamic range of human hearing, dB)

5. Quantization in Practice: Integers vs. Floats

Now we can put this all into the context of your work in AI and machine learning.

  • Storage (Integer Formats): Audio is typically stored in files (like WAV) as 16-bit or 24-bit signed integers. This provides a good balance of fidelity and file size.
  • Processing (Floating-Point Formats): When you load an audio file into a library like torchaudio or librosa, the integer values are almost always converted to 32-bit floating-point numbers, typically scaled to the range [-1.0, 1.0].

Why the conversion? 32-bit floating-point numbers offer a practically immense dynamic range ( dB). This provides massive "headroom," preventing clipping and the accumulation of quantization noise during complex mathematical operations like filtering, mixing, or the forward passes of a neural network. You perform all your processing in this high-precision domain and only convert back to an integer format when saving the final output.

Quantization — Digital Signals Theory

To conclude, let's review a practical summary of how quantization is handled in typical audio processing workflows.

Read sections '2.4.5. Floating point numbers' and '2.4.7. Quantization in practice'. This will solidify the distinction between integer formats for storage and floating-point formats for processing, which is standard practice in audio AI.


Conclusion

In this lesson, we completed the second half of the analog-to-digital conversion process. You have now seen how a continuous wave is transformed into a series of numbers that a computer can store and process.

Key Takeaways:

  • Quantization is the process of mapping a continuous range of amplitudes to a finite set of discrete levels. It is a lossy process that discretizes the amplitude axis.
  • Bit Depth (n) determines the number of available levels () and thus the precision of the amplitude representation.
  • Quantization Error is the rounding error introduced by this process, which manifests as audible quantization noise.
  • Dynamic Range is the primary measure of fidelity affected by bit depth. Each additional bit adds approximately 6 dB of dynamic range, effectively lowering the noise floor.
  • In practice, audio is stored as 16-bit or 24-bit integers but processed as 32-bit floating-point numbers to maintain precision and prevent errors during computation.

We have now defined a digital audio signal by its sampling rate (from the last lesson) and its bit depth. In the next lesson, we will see how these two parameters are used in common digital audio file formats like WAV, FLAC, and MP3, and explore the trade-offs between uncompressed, lossless, and lossy formats.

Can't find a good explanation? Sign up and we'll make it for you

Sign up