Skip to main content
Create your own

Implementing a DCGAN

Hello! Welcome back.

In our previous lesson, we built a basic Generative Adversarial Network using simple multi-layer perceptrons. We saw how the adversarial game between a generator and a discriminator could produce new data, but we also discovered that this process is often unstable and prone to issues like mode collapse.

Today, we'll address those stability issues head-on. Your learning outcome is to implement a Deep Convolutional GAN (DCGAN). A DCGAN isn't a new theoretical framework like the jump from VAEs to GANs; rather, it's a set of powerful architectural guidelines that made GANs practical and effective for image generation tasks.

We will cover:

  • The key architectural principles introduced by the DCGAN paper.
  • The role of convolutional and transposed convolutional layers in the discriminator and generator.
  • How components like Batch Normalization and specific activation functions contribute to training stability.
  • A complete, from-scratch implementation of a DCGAN in PyTorch to generate images.

By the end of this lesson, you'll have built a much more robust and powerful GAN, moving from the simple "vanilla" model to a cornerstone architecture in the history of generative models.

From Vanilla GAN to DCGAN: A Recipe for Stability

The 2015 paper, "Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks," was a landmark publication. It didn't change the fundamental minimax game of GANs, but it provided a concrete recipe for building GANs that train stably and produce high-quality images.

These architectural guidelines are the "secret sauce" of DCGANs. Let's explore them before we start coding.

DCGAN implementation from scratch

The video 'DCGAN implementation from scratch' by Aladdin Persson starts by summarizing the key contributions from the original DCGAN paper. This provides an excellent overview of the architectural rules we'll be following.

Watch the first part of the video (00:28 - 04:16). As you watch, take note of the main architectural guidelines and hyperparameter choices discussed.

To summarize and expand on the key points from the video and the original paper:

  1. Replace Pooling with Strided Convolutions:

    • Discriminator: Instead of using max-pooling or average-pooling layers to downsample the image, DCGANs use strided convolutions (stride > 1). This allows the network to learn its own spatial downsampling, making it more powerful than a fixed function like max-pooling.
    • Generator: To upsample from the latent vector, the generator uses transposed convolutions (often called fractionally-strided convolutions). This is the inverse of a convolution, allowing the network to learn how to intelligently increase spatial dimensions.
  2. Use Batch Normalization: Batch Normalization is applied in both the generator and discriminator. It stabilizes learning by normalizing the input to each layer to have zero mean and unit variance. This helps prevent issues like mode collapse and allows gradients to flow more effectively, which is critical in the delicate GAN training dynamic. There are exceptions: it's not used in the generator's output layer or the discriminator's input layer.

  3. Eliminate Fully Connected Layers: The networks are fully convolutional. This reduces the number of parameters and improves architectural stability. The final layer of the discriminator is flattened and fed to a single sigmoid output, but there are no dense hidden layers.

  4. Choose the Right Activation Functions:

    • Generator: Uses the ReLU activation function for all layers except the output layer. The output layer uses Tanh, which scales the output to the range [-1, 1]. This pairs well with normalizing the input training images to the same range.
    • Discriminator: Uses LeakyReLU for all layers. Unlike a standard ReLU which outputs zero for any negative input, LeakyReLU allows a small, non-zero gradient to flow through for negative values. This prevents "dying ReLUs" and ensures the discriminator's gradients don't vanish, providing a more consistent signal to the generator.
  5. Specific Optimizer Settings: The paper recommends the Adam optimizer with a learning rate of 0.0002 and a beta1 parameter of 0.5 (instead of the default 0.9). This slower, more stable update rule was found to be crucial for GAN convergence.

Let's visualize the architectures we are about to build.

DCGAN Generator Architecture
This diagram shows the DCGAN generator. A 100-dimensional latent vector `z` is first projected and reshaped into a small spatial volume with many channels. A series of transposed convolutions then progressively upsamples this volume, reducing the channel depth and increasing the spatial dimensions until a 64x64x3 image is formed.
DCGAN Discriminator Architecture Diagram
This diagram shows the DCGAN discriminator. It takes a 64x64x3 image and passes it through a series of strided convolutions, which downsample the spatial dimensions while increasing the number of feature maps. The final output is a single scalar value representing the probability that the input image is real.

Implementing a DCGAN in PyTorch

Now, let's translate these architectural guidelines into code. We will follow a detailed, step-by-step implementation. Your background in software engineering will be helpful here as we construct these models as nn.Module classes, paying close attention to the input and output tensor shapes at each layer.

The following sections will guide you through the implementation, using Aladdin Persson's video as a code-along reference.

1. The Discriminator

The discriminator acts as a binary classifier. It takes an image and outputs a single value indicating the probability that the image is real. We'll build it using a sequence of Conv2d, BatchNorm2d, and LeakyReLU layers.

DCGAN implementation from scratch

Let's start by building the discriminator. This section of the video translates the downsampling architecture we saw in the diagram into a PyTorch nn.Sequential model.

Watch from 04:16 to 09:34. Pay attention to: The creation of a reusable block for the Conv-BN-LeakyReLU pattern. The specific kernel size, stride, and padding values used to halve the spatial dimensions at each step. The omission of batch norm in the first layer, as per the paper's guidelines. The final convolution that maps the feature volume to a single output, followed by a Sigmoid activation.

2. The Generator

The generator's job is to do the reverse: take a small latent vector z and upsample it into a full-sized image. This is achieved using ConvTranspose2d layers.

DCGAN implementation from scratch

Now for the generator. This part of the video demonstrates how to use ConvTranspose2d layers to upsample a latent vector into an image.

Watch from 09:34 to 15:17. Focus on: The generator's block, which uses the ConvTranspose2d-BN-ReLU pattern. How the first block reshapes the 1D latent vector into a 4x4 feature map. The subsequent blocks that double the spatial dimensions at each step (4x4 -> 8x8 -> 16x16 -> 32x32 -> 64x64). The final ConvTranspose2d layer followed by a Tanh activation to match the normalized image data range.

Test your understanding!

In the discriminator, we use nn.Conv2d with stride=2 to downsample. In the generator, we use nn.ConvTranspose2d with stride=2 to upsample. Why is ConvTranspose2d a better choice for the generator than, for example, using a simple upsampling method (like nearest-neighbor) followed by a standard nn.Conv2d?

Show answer

While using an upsampling layer followed by a standard convolution is a valid approach (and used in some modern architectures), the ConvTranspose2d layer has a key advantage in the DCGAN context: its weights are learnable. This means the network doesn't just apply a fixed upsampling algorithm (like duplicating pixels); it learns the optimal way to fill in the details and reverse the downsampling process performed by the discriminator's convolutions. This allows the generator to learn how to create more detailed and less "blocky" structures during upsampling.

3. Weight Initialization and Unit Testing

The DCGAN paper specifies a particular weight initialization scheme. It's a small but important detail for stability. After defining our models, it's also good practice to run a quick sanity check to ensure our tensor shapes are correct.

DCGAN implementation from scratch

Let's implement the weight initialization function and run a quick test on our models.

Watch from 15:17 to 18:54. This short segment covers: A function initialize_weights that sets the weights of all convolutional and batch norm layers from a normal distribution with mean=0 and stdev=0.02. A simple test function that passes random noise through the generator and a random image through the discriminator to assert that the output shapes are correct. This is a great debugging practice.

4. The Training Loop

The training logic for a DCGAN is identical to the vanilla GAN we saw in the previous lesson. We'll alternate between training the discriminator and the generator.

  • Part 1: Train the Discriminator
    1. Feed it a batch of real images and calculate the loss against labels of 1.
    2. Feed it a batch of fake images (from the generator) and calculate the loss against labels of 0.
    3. Sum the two losses and update the discriminator's weights.
  • Part 2: Train the Generator
    1. Feed a batch of fake images to the discriminator.
    2. Calculate the generator's loss based on the discriminator's output, but this time using labels of 1 (the generator's goal is to make the discriminator think its fakes are real).
    3. Update the generator's weights.

The following video sections set up the hyperparameters and implement this training loop.

DCGAN implementation from scratch

With our models defined, we can now set up the training script and implement the core adversarial training loop.

Watch from 19:02 to 31:33. This is the final and most crucial part of the implementation. Training Setup (19:02 - 24:35): Observe the setup of hyperparameters (learning rate, beta1, etc.), data transformations (including normalizing images to [-1, 1]), and the Adam optimizers. Training Loop (24:35 - 31:33): Follow the implementation of the two-part training process. Note how errD_real and errD_fake are calculated and combined to update the discriminator, and how errG is calculated to update the generator.

After running the code, you'll be able to generate images of celebrities (if using the CelebA dataset) or other objects, which are of significantly higher quality and stability than what a simple MLP-based GAN could produce.

Bonus: An Alternative Perspective (Keras & Anime Faces)

Given your interest in anime, you might find this alternative implementation interesting. It uses Keras/TensorFlow to train a DCGAN on a dataset of anime faces. While the core principles are the same, the implementation details offer a different perspective.

Anime Face Generation using DCGAN | Keras Tensorflow | Deep Learning | Python

This video from Hackers Realm builds a DCGAN for anime face generation using Keras. It's a great example of applying these techniques to a different domain and framework.

This is optional, but you might find it insightful. You can skim through these sections: Generator (19:32 - 29:47): See how a similar architecture is built using the Keras Sequential API. Discriminator (29:47 - 37:33): Notice the same pattern of strided convolutions and LeakyReLU. Custom train_step (37:33 - 01:02:37): Keras often uses a custom Model class that overrides the train_step method. This encapsulates the entire adversarial update logic within the model itself, which is a different style from the explicit PyTorch loop. Debugging (01:14:33 - 01:20:12): The creator runs into a common issue (exploding loss) and debugs it by removing a BatchNorm layer from the generator. This is a valuable, real-world lesson in how sensitive GANs can be.

Conclusion

Congratulations! You've just implemented a Deep Convolutional GAN, a foundational architecture in generative modeling. We've moved beyond the toy examples of fully-connected GANs and built a model capable of learning from complex image datasets.

Key Takeaways:

  • DCGANs are not a new theory but a set of architectural best practices for building stable GANs.
  • Strided and transposed convolutions allow the networks to learn their own spatial upsampling and downsampling.
  • Batch Normalization is a key stabilizer, helping to regulate gradient flow and prevent mode collapse.
  • Specific activation functions (LeakyReLU in the discriminator, Tanh output in the generator) and optimizer settings are crucial for successful training.
  • The implementation requires careful management of tensor shapes as data flows through the convolutional layers of the generator and discriminator.

Preview of the next lesson:
The DCGAN architecture significantly improves training stability. However, the underlying loss function (Binary Cross-Entropy) still has issues. It can saturate, cause vanishing gradients, and doesn't correlate well with the visual quality of the generated images. In our next lesson, we will explore Wasserstein GANs (WGANs), which introduce a new, theoretically grounded loss function to address these problems, leading to even more stable training and a meaningful loss metric.

Can't find a good explanation? Sign up and we'll make it for you

Sign up