Hello! Welcome to your next lesson on generative models.
In our last session, we made a significant leap by implementing Wasserstein GANs with Gradient Penalty (WGAN-GP). We moved beyond the standard GAN loss function to a more stable objective based on the Wasserstein distance. This gave us a reliable way to train GANs, where the critic's loss actually correlates with image quality, largely solving the notorious problems of training instability and mode collapse.
Now that we have a robust training framework, we can build more powerful and sophisticated architectures upon it. Today's lesson is about pushing the boundaries of what GANs can do in terms of both control and quality.
Your learning outcome is to implement advanced GAN architectures for controlled and high-fidelity generation (Conditional GANs, StyleGAN).
We will explore two major architectural advancements:
- Conditional GANs (cGANs): We'll learn how to direct the GAN's creative process, telling it what to generate, for instance, by providing a class label or even another image as a condition.
- StyleGAN: We'll dissect the architecture that powered a revolution in photorealistic image generation, famous for its ability to create stunningly high-quality images and offer unprecedented, intuitive control over the visual style.
By the end of this lesson, you'll understand the principles behind these advanced models and how to implement them.
1. Gaining Control: Conditional GANs (cGANs)
So far, our GANs have been unconditional. We train them on a dataset (e.g., celebrity faces), and they learn to generate new samples from that same distribution. But we have no control over the output. If we want a specific type of face—say, one with glasses—we're out of luck.
Conditional GANs solve this by introducing an extra piece of information, a condition y, into the model. This y can be anything: a class label, a text description, or even a whole other image.
The core idea is simple but effective: feed the condition y to both the generator and the discriminator.

This changes the objective for each network:
- Generator: Must now generate an image
xthat is not only realistic but also matches the given conditiony. - Discriminator: Must now determine if an image
xis a real sample from the dataset and if it is paired with the correct conditiony. A real image with the wrong label is considered "fake" by the discriminator.
Conditional GAN (cGAN) in PyTorch and TensorFlow
To understand the roles of the conditioned generator and discriminator in more detail, let's read the introductory sections of the article 'Conditional GAN (cGAN) in PyTorch and TensorFlow' from LearnOpenCV.
Read the sections 'What is a Conditional GAN?' and 'Purpose of Conditional Generator and Discriminator'. This will clarify how the responsibilities of both networks are expanded to handle conditional information.
Implementing Conditioning
In practice, how do we feed this condition y into the network? A common technique for categorical labels (like MNIST digits or class labels from a dataset) is to use an Embedding layer.
- The integer class label is passed to an
nn.Embeddinglayer, which converts it into a dense vector. - This vector is then reshaped and concatenated with the latent vector
zin the generator, or with the feature maps in the discriminator, at an appropriate layer.
A Powerful cGAN: Pix2Pix for Image-to-Image Translation
One of the most compelling applications of cGANs is image-to-image translation, where the "condition" is an entire input image. The goal is to learn a mapping from an input image to a corresponding output image. Examples include:
- Converting satellite photos to maps.
- Turning architectural sketches into photorealistic buildings.
- Colorizing black-and-white photos.
The Pix2Pix model is a seminal architecture for this task. It introduces two key components:
- U-Net Generator: Instead of a standard encoder-decoder, it uses a U-Net architecture. The "skip connections" in the U-Net pass low-level information (like edges and textures) directly from the encoder to the decoder. This is crucial for tasks where the output structure must align with the input.
- PatchGAN Discriminator: Instead of classifying the entire image as real or fake, the PatchGAN discriminator looks at small
N x Npatches of the image and classifies each patch. This encourages local, high-frequency realism and is computationally more efficient.
The video below provides an excellent walkthrough of the Pix2Pix architecture and its implementation. Since you're already familiar with the author's clear, code-first style from our previous lessons, this will be a great way to see a cGAN built from scratch.
Pix2Pix implementation from scratch
Let's watch Aladdin Persson's implementation of Pix2Pix. He clearly explains the motivation and then dives into building the U-Net generator and PatchGAN discriminator in PyTorch.
Watch the following segments to understand and implement the core components of Pix2Pix: Paper Summary (00:11 - 05:48): A quick overview of the Pix2Pix paper, explaining the motivation for using a GAN, the U-Net generator, and the PatchGAN discriminator. Discriminator Implementation (09:08 - 20:15): A step-by-step implementation of the PatchGAN discriminator. Pay attention to how it takes two images (input and target/generated) and outputs a patch of predictions. Generator Implementation (20:15 - 34:27): A detailed implementation of the U-Net-style generator. Focus on how the downsampling and upsampling blocks are constructed and how the skip connections are concatenated. Training Setup & Loop (43:41 - 57:39): This section covers setting up the optimizers, loss functions (including the L1 loss which encourages the generator to be close to the ground truth), and the main training steps for the discriminator and generator. The logic here directly builds on the GAN training loops you've already implemented.
Test your understanding!
The Pix2Pix generator uses a U-Net architecture with skip connections. What problem would you expect to see in the generated images if these skip connections were removed, leaving a standard encoder-decoder architecture?
Show answer
Without skip connections, the low-level information from the input image (like precise edges, textures, and structural details) has to pass through the entire network, including the bottleneck. This information compression often leads to blurry and less detailed outputs. The skip connections provide a direct "shortcut" for this high-frequency information, allowing the generator to reconstruct sharp, detailed images that are structurally consistent with the input.
2. Pushing for Quality: StyleGAN
While cGANs give us control, a separate line of research focused on a different question: how can we generate images of unprecedented quality and resolution? The StyleGAN family of models represents a monumental leap in this direction.
StyleGAN introduces a completely redesigned generator that moves away from the traditional DCGAN approach. It's built on several key innovations.
Key Architectural Ideas of StyleGAN
-
Progressive Growing (from ProGAN): Training a model to directly output high-resolution images (e.g., 1024x1024) from the start is very unstable. StyleGAN borrows the core idea from its predecessor, ProGAN, to train progressively. The model starts by generating tiny 4x4 images, and once that is stable, new layers are faded in to double the resolution (8x8, 16x16, and so on). This allows the network to learn coarse, high-level features first before moving on to fine details, dramatically stabilizing training.
-
Mapping Network (Z → W): Instead of feeding the latent vector
zdirectly into the synthesis network, StyleGAN first passes it through a non-linear mapping network (an 8-layer MLP). This produces an intermediate latent vectorwin a new spaceW. The purpose of this is to "disentangle" the factors of variation in the input latents. TheWspace is less correlated, meaning that changing one dimension ofwtends to correspond to a single, clean semantic change in the final image (like changing age, or smiling), which is difficult withz. -
Style-Based Generator and AdaIN: The most significant innovation is that the intermediate latent
wis not fed into the beginning of the synthesis network. Instead, it is used to control the "style" at each resolution level.wis transformed and fed into Adaptive Instance Normalization (AdaIN) layers after each convolution. The AdaIN layer first normalizes the feature map activations and then applies a learned scale and bias derived fromw. This allowswto control the visual features at different scales, from coarse styles (pose, face shape) at low resolutions to fine styles (hair texture, skin color) at high resolutions. -
Stochastic Variation via Noise Injection: To model fine, stochastic details like freckles, pores, or individual hair strands, explicit noise is added to the feature maps at each resolution. This is a brilliant move: it separates the high-level style (controlled by
w) from the random, per-pixel details, which are controlled by the noise maps. Generating the same image with different noise maps will result in the same person but with slightly different hair placement or skin texture.
This video gives a great theoretical overview and code breakdown of StyleGAN, starting with the ProGAN concepts it builds upon.
StyleGAN 1 Guide [Theory and PyTorch Code, ProGAN included]
To understand these revolutionary concepts, let's watch the 'StyleGAN 1 Guide' by Moran Reznik. The video explains the transition from ProGAN to StyleGAN and the key components like AdaIN.
Watch the following segments to grasp the core ideas of StyleGAN: ProGAN Introduction (01:15 - 04:45): Understand the concept of progressive growing, which is the foundation of StyleGAN's training process. StyleGAN's Main Changes (04:45 - 05:24): A concise summary of the two major architectural shifts: moving the latent code to normalization functions and adding stochastic noise. Adaptive Instance Normalization (AdaIN) (06:53 - 09:14): This is the heart of StyleGAN. Pay close attention to how the latent code w is used to generate the mean and standard deviation to control the feature maps. Results and Style Control (10:08 - 11:31): See the power of this architecture in action. Observe how manipulating the style vector w at different resolutions affects different attributes of the generated face.
StyleGAN2: A More Refined Implementation
StyleGAN2 improved upon the original by simplifying the architecture and removing normalization artifacts. For a deeper dive into a production-quality implementation, the following article provides a complete from-scratch breakdown in PyTorch. Your CS background will make the modular code structure easy to follow.
Implementation StyleGAN2 from scratch
For an in-depth look at a more modern implementation, the article 'Implementation StyleGAN2 from scratch' on Paperspace provides a detailed code walkthrough. It's an excellent resource for understanding how these complex components fit together.
Review the code snippets for these key components to see how the theory is translated into PyTorch modules: Noise Mapping Network: This implements the z to w mapping. Generator: Note how it's composed of multiple GeneratorBlock modules and how the RGB outputs are progressively summed. Style Block: This is the core synthesis block, containing the convolution and style modulation. Convolution with Weight Modulation and Demodulation: This is the improved mechanism in StyleGAN2 that replaces AdaIN. It directly scales the convolution weights based on the style vector w.
Test your understanding!
In the StyleGAN architecture, what is the fundamental difference in the roles of the intermediate latent vector w and the injected noise maps?
Show answer
The intermediate latent vector w controls the high-level, global "style" of the image. By modulating the convolutional layers via AdaIN (or weight modulation in StyleGAN2), it dictates semantic attributes like identity, pose, face shape, and color schemes. In contrast, the noise maps provide stochastic, localized variation. They are responsible for fine, non-semantic details like the exact placement of hairs, freckles, or skin pores. This separation allows for independent control over the core identity/style and the random details of an image.
Conclusion
Excellent work! Today, we've explored two major paths for advancing GAN architectures beyond the basics. You've seen how to take a stable training framework like WGAN-GP and build upon it to achieve greater control and photorealistic quality.
Key Takeaways:
- Conditional GANs (cGANs) provide explicit control over the generation process by feeding conditional information (like class labels or images) to both the generator and discriminator.
- Pix2Pix is a powerful cGAN for image-to-image translation that uses a U-Net generator with skip connections to preserve structure and a PatchGAN discriminator to enforce local realism.
- StyleGAN achieves state-of-the-art image quality through a radically new generator architecture featuring:
- Progressive growing for stable high-resolution training.
- A mapping network that creates a disentangled intermediate latent space
W. - Style modulation via AdaIN, where
wcontrols features at every scale. - Stochastic noise injection to model fine, random details independently of the style.
Preview of the next lesson:
While GANs have been the dominant force in generative modeling for years, a different paradigm has recently risen to prominence, often surpassing GANs in image quality and training stability. In the next lesson, we will begin our journey into this new and exciting area by exploring Denoising Diffusion Probabilistic Models (DDPMs). We will implement the forward (noising) and reverse (denoising) processes that form the foundation of models like DALL-E 2 and Stable Diffusion.