Hello! Welcome to the next lesson in our journey into audio AI.
In our last session, we successfully containerized our speech recognition API using Docker. This was a massive step forward, solving the problem of reproducing our code and its environment anywhere. However, an ML application has more than just code. It has large datasets, multi-gigabyte model weights, and various other artifacts. A git repository is designed for code, not for storing a 5GB Whisper model. Pushing such files to GitHub would be slow, inefficient, and quickly exceed repository size limits.
This brings us to a crucial question in MLOps: How do we version our data and models in sync with our code, ensuring that our entire experiment—code, data, and model—is fully reproducible?
Today, we will address this by focusing on the learning outcome: Manage model artifacts and data versions for reproducibility using a tool like DVC. We'll learn how to use Data Version Control (DVC) to handle large files, allowing us to version our entire ML project seamlessly.
1. The DVC Philosophy: Git for Code, DVC for Data
At its core, DVC is not a replacement for Git. Instead, it works alongside Git to provide a complete versioning solution. The philosophy is simple:
- Git handles what it's good at: versioning text-based source code and small configuration files.
- DVC handles what Git is bad at: versioning large binary files like datasets and model weights.
So, how does it work? When you tell DVC to track a large file (e.g., model.pt), it does two things:
- It moves the actual file to a hidden cache and tells Git to ignore it (by adding it to
.gitignore). - It creates a tiny text file,
model.pt.dvc, which acts as a pointer. This pointer file contains metadata, like a hash of the original file, that uniquely identifies it.
This small .dvc file is then committed to your Git repository. The large file itself is pushed to a separate remote storage location (like Google Drive, AWS S3, or even a folder on a server).
Think of it like a library catalog. Your Git repository is the card catalog, full of small, easy-to-manage cards (.dvc files). Each card tells you exactly where to find the huge, heavy book (your data file) in the library's main storage.

2. The Core DVC Workflow: A Hands-On Tutorial
The best way to understand DVC is to use it. We'll now walk through the essential commands for versioning a file. Imagine you have your Whisper model file that you want to track.
The following video from the creators of DVC provides an excellent, concise demonstration of the entire basic workflow. We will watch it in segments, and I will provide commentary on each part.
Versioning Data with DVC (Hands-On Tutorial!)
This hands-on tutorial from DVC.org will guide us through initializing a DVC project, tracking a file, setting up remote storage, and switching between different versions of our data.
Please watch the video from the beginning until 10:09. I've broken it down into three key phases below. We'll pause after each to discuss what happened.
Let's break down the process you just saw.
Phase 1: Initialization and Tracking a File (00:20 - 02:39)
First, you need to set up your project to use DVC. In your project's root directory (which should already be a Git repository), you run:
# Initialize DVC in the current git repository
dvc init
This creates a .dvc/ directory, which DVC uses to store its configuration and local cache.
Next, to start tracking a large file or directory, you use dvc add. Let's assume you have your model in models/whisper.pt.
# Tell DVC to start tracking the model file
dvc add models/whisper.pt
As the video explains, this command does a few things:
- It creates a pointer file:
models/whisper.pt.dvc. - It adds
models/whisper.ptto a.gitignorefile so Git will ignore the large file. - It copies the actual
whisper.ptfile into DVC's local cache (.dvc/cache).
Now, you commit the pointer file to Git. This saves a "version" of your data that corresponds to this point in your code's history.
# Add the new .dvc file and the updated .gitignore to git's staging area
git add models/whisper.pt.dvc .gitignore
# Commit the pointer file
git commit -m "Track initial Whisper model"
Phase 2: Remote Storage and Pushing (02:39 - 05:46)
Your model is now versioned locally, but the large file only exists on your machine. To share it, you need to configure remote storage. DVC supports many backends (S3, GCS, Azure, SSH), but a simple option to start with is Google Drive, as shown in the video.
# Configure a default remote storage named 'gdrive_storage'
# You would replace the ID with the one from your own Google Drive folder
dvc remote add -d gdrive_storage gdrive://<your_gdrive_folder_id>
# Commit the updated DVC config to Git
git add .dvc/config
git commit -m "Configure DVC remote storage"
With the remote configured, you can now push your data. This is a two-step process:
git push: Pushes your code and the small.dvcpointer files to your Git remote (e.g., GitHub).dvc push: Pushes the actual large model file from your local cache to your configured DVC remote storage (e.g., Google Drive).
# Push the large file to your Google Drive remote
dvc push
Phase 3: Pulling and Switching Versions (05:46 - 10:09)
This is where the magic happens. Imagine a colleague (or you, on a different machine) clones your project:
git clone <your_github_repo_url>
cd <your_project>
They now have the code and the models/whisper.pt.dvc pointer file, but not the large model itself. To get the model, they simply run:
# Download the data from the remote storage into the workspace
dvc pull
dvc pull inspects the .dvc file, finds the corresponding data in the remote storage, downloads it to the local cache, and links it into the workspace.
Now, let's see how to switch between versions. Suppose you fine-tuned your model and created a new version. You would dvc add it and git commit the changes, creating a new version history.
To revert to a previous version of your model, you use Git to check out the old pointer file, and then DVC to sync your workspace.
# 1. Use git to check out the previous commit where you tracked the v1 model.
# This will revert the models/whisper.pt.dvc pointer file to its old state.
git checkout <commit_hash_for_v1> models/whisper.pt.dvc
# 2. Tell DVC to sync your workspace with the state of the pointer files.
# This will replace the v2 model file with the v1 model file from the cache.
dvc checkout
This workflow powerfully links your data versions directly to your Git commit history, achieving true reproducibility.
3. Beyond Simple Tracking: Reproducible Pipelines
While dvc add is great for versioning static assets, much of our work involves generating artifacts from code (e.g., training a model from a dataset). DVC pipelines allow you to codify these steps, making your entire workflow reproducible.
A DVC pipeline stage defines:
- Dependencies (
-d): Inputs, like source code (train.py) and data (data/). - Outputs (
-o): Artifacts produced, like a model file (model.pt). - Metrics (
-M): JSON/YAML files with performance metrics (metrics.json). - Command: The shell command to execute (
python train.py).
DVC stores this pipeline definition in a human-readable dvc.yaml file. The real power comes from the dvc repro command, which intelligently re-runs stages only if their dependencies have changed.
Data Version Control With Python and DVC
This article from Real Python provides an excellent tutorial on building DVC pipelines. We'll focus on the section that transforms a series of manual steps into an automated, reproducible pipeline.
Please read the section titled 'Create Reproducible Pipelines.' Pay close attention to: The structure of the dvc run command, with its -n (name), -d (dependencies), and -o (outputs) flags. How multiple dvc run commands chain together to form a pipeline, with the output of one stage becoming the dependency of the next. The concept of the dvc.yaml and dvc.lock files that are created to store the pipeline definition. The use of dvc repro, which automatically reruns the necessary parts of the pipeline when a dependency like a source file is modified.
By defining your data processing, training, and evaluation steps in a dvc.yaml file, you create a DAG (Directed Acyclic Graph) of your entire ML workflow. This is a cornerstone of MLOps, ensuring that anyone can reproduce your results with a single dvc repro command.
4. DVC and Docker: A Perfect Match for Deployment
So, how does DVC fit in with the Docker container we built in the last lesson? They work together to ensure a fully reproducible deployment.
The key is that you do not run dvc pull inside your Dockerfile. Instead, your Continuous Integration/Continuous Deployment (CI/CD) pipeline performs these steps in order:
- Check out Code:
git clone my-repo.git && cd my-repo - Pull Data/Model:
dvc pull(This uses the.dvcfiles from Git to download the correct large model artifact into the local workspace). - Build Container:
docker build -t my-speech-service .
Your Dockerfile can now be simplified. Instead of downloading a model from the internet, it just copies the model that dvc pull placed in the local directory.
# ... (other Dockerfile instructions)
# DVC has already pulled the model into the './models' directory
# Now we just copy it into the image.
COPY ./models /app/models
# ... (rest of the Dockerfile)
This makes your Docker builds faster, more reliable, and completely decoupled from external services like Hugging Face Hub. The specific model version being packaged is determined by the Git commit you are building from, ensuring end-to-end version control.
Conclusion
Today you've learned how to solve a fundamental problem in MLOps: versioning large data and model artifacts. By combining Git's strength in managing code with DVC's ability to handle large files, you can create fully reproducible machine learning projects.
Key Takeaways:
- DVC works with Git to version large files by storing them in remote storage and keeping small "pointer" files in the Git repository.
- The core workflow involves
dvc addto track files,dvc pushto upload them, anddvc pullto download them. - You can switch between data/model versions by using
git checkouton the.dvcfile, followed bydvc checkout. - For reproducible workflows,
dvc runanddvc reproallow you to define and execute multi-stage pipelines that only re-run when inputs change. - In a CI/CD context, you run
dvc pullbeforedocker buildto ensure the correct, versioned model is packaged into your container.
Preview of the Next Lesson:
We now have the tools to create isolated, reproducible services (Docker) and version all the artifacts that go into them (DVC). We're ready to think bigger. In the next lesson, we will design the system architecture for a conversational AI agent, exploring how to combine our version-controlled STT and TTS components into a robust, scalable pipeline.