Hello! Welcome to the first lesson of your course on Audio AI.
Your goal is to become an audio researcher and developer, and you've asked for a comprehensive path that starts from the absolute fundamentals. That's exactly where we'll begin. Over this course, we'll build your understanding from the ground up, moving from the physics of sound to the sophisticated neural networks used in modern speech and audio systems.
This first lesson addresses a foundational learning outcome: modeling a sound wave mathematically. We will explore how the physical phenomenon of a sound wave can be described by a simple but powerful mathematical equation, defined by three core properties: amplitude, frequency, and phase. Mastering this is the essential first step before we can digitize, process, and ultimately generate audio with AI.
1. What is Sound? From Physics to Waveforms
At its core, sound is a mechanical wave—a vibration that propagates through a medium like air. When a source like a speaker vibrates, it pushes and pulls on the air molecules around it, creating alternating regions of high pressure (compressions) and low pressure (rarefactions). This pressure disturbance travels outwards as a longitudinal wave.

The simplest and most fundamental type of sound wave is a pure tone, which can be perfectly described by a sinusoidal wave (a sine or cosine curve).
To understand why sinusoids are so important, read the short section linked below from the resource Mathematics of the DFT.
Sinusoids and Exponentials | Mathematics of the DFT
This section explains why sinusoids are fundamental not only in physics (as in the simple harmonic motion of a tuning fork) but also because the human ear itself acts as a biological spectrum analyzer, breaking down complex sounds into sinusoidal components.
Please read the section titled 'Why Sinusoids are Important'.
This idea that the ear perceives sound in terms of its frequency components is a cornerstone of audio processing. It's why modeling sound with sinusoids is not just a mathematical convenience, but a perceptually relevant approach.
2. The Anatomy of a Sound Wave
Since a pure tone can be represented by a sine wave, we can use the mathematical properties of a sinusoid to describe the sound. Let's break down the key parameters.

The general mathematical model for a sinusoidal wave as a function of time is:
Let's dissect each component of this equation. For a more detailed walkthrough, the following video provides an excellent explanation.
This video from UNSW Physics provides a clear, step-by-step breakdown of the sinusoidal wave equation. It defines each variable and works through a practical example of how to extract the wave's properties from its equation.
Please watch from 01:00 to 07:24. The first part (01:00-02:18) defines the components of the equation. The rest of the video (starting at 03:28) works through a concrete example problem, which is very helpful for solidification.
Now, let's formally define each term in the context of an audio signal.
Amplitude (A)
- Definition: The amplitude is the maximum displacement or distance moved by a point on a vibrating body or wave measured from its equilibrium position. In our equation, this is represented by .
- Perceptual Correlate: Loudness. A larger amplitude corresponds to a more intense pressure variation, which our ears perceive as a louder sound.
- Unit: In physics, it can be pressure (Pascals). In a digital system, it's often represented by a normalized value, typically between -1.0 and 1.0.
Frequency (f)
- Definition: The frequency is the number of complete oscillations or cycles that occur per unit of time. In our equation, this is represented by .
- Perceptual Correlate: Pitch. A higher frequency corresponds to a higher-pitched sound. For example, the musical note "A4" (the A above middle C) has a standard frequency of 440 Hz.
- Unit: Hertz (Hz), where 1 Hz = 1 cycle per second.
Two other closely related terms are:
- Period (T): The time it takes to complete one full cycle. It is the reciprocal of frequency: .
- Angular Frequency (): This is a measure of rotation rate, expressed in radians per second. It's used because it simplifies the sinusoidal equation. The relationship is . Using angular frequency, our equation becomes: This form is very common in engineering and signal processing.
Phase ()
- Definition: The phase, or phase offset, describes the starting position of the sine wave at time . It represents a horizontal shift of the wave. In our equation, this is represented by .
- Perceptual Correlate: The human ear is not very sensitive to the absolute phase of a single, static sound wave. However, the relative phase between multiple waves is critical. When waves combine, their phases determine whether they reinforce each other (constructive interference) or cancel each other out (destructive interference). This is fundamental to the timbre and texture of complex sounds.
- Unit: Radians or degrees. By convention in signal processing, we almost always use radians.
To solidify these definitions, please read the following brief sections.
Real Python: Reading and Writing WAV Files & Mathematics of the DFT
These resources provide concise, formal definitions of the terms we've just discussed.
From the Real Python article (resource_id 'LINK'), read the section 'The Waveform Part of WAV'. Focus on the definitions of Amplitude, Frequency, and Phase.
Sinusoids and Exponentials | Mathematics of the DFT
From the 'Mathematics of the DFT' article (resource_id 'LINK'), read the section titled 'Sinusoids'. This provides the formal mathematical definition and introduces the term 'radian frequency'.
3. From Pure Tones to Complex Sounds
So far, we've only discussed pure tones. Real-world sounds, like your voice, a piano note, or a drum hit, are far more complex. The remarkable insight of Fourier's theorem is that any complex periodic sound can be broken down into a sum of simple sinusoids, each with its own amplitude, frequency, and phase.
The base frequency is called the fundamental, and it determines the overall pitch we perceive. The additional sinusoids, called overtones or harmonics (which are integer multiples of the fundamental), are what give an instrument its unique character or timbre.
The following video provides a great visual and auditory demonstration of this principle.
The Math Behind Music and Sound Synthesis
This video from Gonkee explains Fourier's theorem in the context of sound synthesis. It demonstrates how complex, rich-sounding waves (like square and saw waves) can be constructed by adding together simple sine waves.
Watch the segment from 05:22 to 05:56, which introduces Fourier's theorem for sound. Then, watch from 07:37 to 10:21, which shows how basic synthesizer waveforms are built from summing sine waves.
This concept is the absolute foundation of spectral analysis, which we will dive into deeply in Module 2. For now, the key takeaway is that our simple sinusoidal model, , is not just a model for pure tones; it is the fundamental building block for all complex sounds.
Conclusion
In this lesson, we established the fundamental link between the physics of sound and its mathematical representation.
Key Takeaways:
- Sound is a pressure wave that can be modeled mathematically.
- The simplest sound, a pure tone, is represented by a sinusoidal wave.
- A sinusoid is completely described by three parameters:
- Amplitude (A): Corresponds to perceived loudness.
- Frequency (f): Corresponds to perceived pitch.
- Phase (): The starting point of the wave, critical for how waves combine.
- The governing equation for a sound wave is or .
- Complex sounds are sums of many sinusoids, making this simple model the universal building block for audio.
You've now built the first, most critical piece of your foundation. You understand the continuous, analog nature of sound. In our next lesson, we will tackle the question: How do we convert this continuous wave into a series of numbers that a computer can store and process? This will lead us to the concepts of sampling, the Nyquist-Shannon theorem, and the problem of aliasing.