Skip to main content
Create your own
Lesson illustration

ElevenLabs Voiceover Tutorial

Hello! Welcome back to your course on building and scaling SaaS companies.

In our last lesson, you mastered a crucial workflow: transforming your video blueprint into a polished, persuasive script using ChatGPT and the powerful PAS (Problem-Agitate-Solve) framework. You now have the written narrative for your product demo video. The words are ready, but for video, the delivery is just as important as the message itself.

Today, we transition from text to audio. Our goal is to generate a natural-sounding voiceover for your video script using ElevenLabs. This is a critical step in your journey to rapidly scale content creation. Instead of spending time and money on recording studios or voice actors, you'll learn how to produce broadcast-quality audio directly from your script, giving you incredible leverage for your marketing efforts.

Introduction to ElevenLabs: Your AI Voice Actor

ElevenLabs is a leading AI speech synthesis platform that goes far beyond simple text-to-speech (TTS). It doesn't just read words; its more advanced models understand the context, emotion, and pacing of a script, allowing you to "direct" an AI voice actor to produce a truly human-like performance.

For a quick overview of what ElevenLabs can do and how to get started, this guide is an excellent entry point.

How to Use ElevenLabs - Best Text to Speech AI Voices (FULL GUIDE)

The video 'How to Use ElevenLabs' by Alec Wilcock provides a fantastic introduction to the platform's core features. It covers everything from the basic interface to the different voice models and settings.

Please watch the first three and a half minutes of this video: Introduction & Overview (00:00 - 01:53): This section explains what ElevenLabs is, its core capabilities, and the pricing structure. Note the mention of a commercial license, which is crucial for your SaaS business goals. Speech Synthesis Interface (01:53 - 03:37): Pay close attention to the user interface. You'll see the main text box, the voice selection menu, and the settings panel. Notice how the pre-made voices are tagged by accent, tone, and use case.

The Core Workflow: From Script to Speech

Now that you're familiar with the interface, let's walk through the process of generating your voiceover. We'll focus on the "Speech Synthesis" tool, which is where you'll spend most of your time.

ElevenLabs Text-to-Speech Interface
This is the main "Speech Synthesis" workspace in ElevenLabs. You'll paste your script into the text box, choose a voice and model, adjust the settings, and generate your audio.

1. Choose Your Voice and Model

Your first decision is selecting an "actor" (the voice) and a "brain" (the AI model).

  • Voice Library: ElevenLabs offers a vast library of pre-made voices. You can filter by gender, accent, and style (e.g., narration, conversational). Take a few minutes to preview different voices to find one that matches the brand you want to build.
  • AI Model: This is the most important setting. ElevenLabs offers several models, but the two you'll use most are:
    • Eleven Multilingual v2: A stable and highly realistic model. Its delivery is primarily controlled by sliders for Stability and Clarity.
    • Eleven v3 (Alpha/Expressive): The newest and most powerful model. It's designed for highly emotional and expressive speech. Instead of sliders, you primarily control it using descriptive "audio tags" directly in your script.

For creating compelling marketing videos, we will focus on the Eleven v3 model due to its superior ability to convey emotion and natural intonation.

2. Fine-Tuning Your Voiceover with the v3 Model

The real power of ElevenLabs comes from your ability to direct the AI's performance. With the v3 model, this is done primarily through text-based commands, which should feel intuitive given your computer science background. Think of them as special syntax for controlling the output.

The official ElevenLabs channel has an excellent tutorial that focuses specifically on this advanced model.

How to make AI Voiceovers that sound Human (2025 ElevenLabs Text to Speech Tutorial)

Watch this tutorial from the ElevenLabs team, 'How to make AI Voiceovers that sound Human'. It dives deep into the expressive capabilities of the v3 model and the use of audio tags.

Focus on the section from 03:12 to 04:52. This part explains and demonstrates how to use audio tags—like [laughs] or [happily]—to direct the v3 model's delivery. This is the key technique for creating a natural-sounding voiceover.

As you saw, audio tags give you director-level control. Here are the key techniques you'll use to prepare your script for the v3 model:

  • Audio Tags for Emotion and Actions: Enclose a description of the desired tone or a non-verbal sound in square brackets []. The AI will interpret this as a direction, not as text to be spoken.

    • Emotion: [enthusiastic], [thoughtful], [sarcastic], [curious]
    • Actions: [laughs], [sighs], [clears throat], [happy gasp]
  • Punctuation for Pacing: The v3 model is highly sensitive to punctuation.

    • Ellipses (…) create a natural, thoughtful pause.
    • An em dash (—) can create a short break.
    • Capitalization (ALL CAPS) adds emphasis or volume to a word or phrase.
    • Exclamation points (!) naturally increase the energy of the delivery.
  • Stability Slider: Even with the v3 model, the Stability slider is important.

    • Low Stability (Creative): More emotional and expressive, but can sometimes lead to unexpected results ("hallucinations"). Best for short, dynamic lines.
    • Medium Stability (Natural): A good balance that's close to the original voice but still responsive to your tags. This is a great starting point.
    • High Stability (Robust): Very consistent and predictable, but less responsive to your emotional tags. It behaves more like the older v2 model.

For a comprehensive list of best practices and tags, the official ElevenLabs documentation is the ultimate resource.

Best practices | ElevenLabs Documentation

The 'Best practices' guide from ElevenLabs' own documentation is your reference manual for fine-tuning. We'll focus on the section dedicated to the v3 model.

Skim the section titled 'Prompting Eleven v3 (alpha)'. You don't need to memorize it, but pay attention to the variety of tags available under 'Audio tags' (Voice-related, Sound effects) and the examples of how punctuation affects delivery. Bookmark this for future reference.

Practical Application: Voicing Your ScribeAI Script

Let's apply this to the "ScribeAI" script you developed in the last lesson. Here is a snippet from the Problem and Agitate sections:

Original Script (Voiceover): "Tired of writing onboarding documents? You spend hours taking screenshots, writing out every single step, and formatting it all, only for the software to update a week later. It's tedious, and it keeps your team from doing their actual jobs."

Now, let's enhance this script with v3 tags to create a more natural and engaging delivery.

Enhanced Script for Eleven v3:
"[sighs] Tired of writing onboarding documents? You spend HOURS taking screenshots… writing out every single step… and formatting it all… only for the software to update a week later. [frustrated] It's tedious, and it keeps your team from doing what they were actually hired to do."

Analysis of Changes:

  • [sighs] adds an immediate, relatable feeling of exhaustion.
  • HOURS in all caps adds emphasis.
  • … (ellipses) create pauses, mimicking someone listing off a series of frustrating tasks.
  • [frustrated] explicitly sets the tone for the final sentence.
  • The wording is slightly changed to sound more conversational ("actual jobs" -> "what they were actually hired to do").
Test your understanding!

Here is the "Solve" part of our ScribeAI script:

"Meet ScribeAI. Just record your screen, and ScribeAI automatically creates a perfect, step-by-step guide with text and annotated screenshots. Now you can export it to Confluence with one click. Your documentation is done in seconds."

How would you modify this script using audio tags and punctuation for the Eleven v3 model to convey a tone of excitement, relief, and confidence?

Show answer

Here’s one possible way to enhance it:

"[confident] Meet ScribeAI. Just hit record on your screen… and ScribeAI automatically creates a perfect, step-by-step guide — complete with text and annotated screenshots. And now… [excited] you can export it to Confluence with ONE click! [triumphant] Your documentation is done. In seconds."

Enhancements:

  • [confident] sets an assured tone from the start.
  • The em dash (—) and pauses (…) build anticipation.
  • [excited] and ONE in caps highlight the key feature.
  • [triumphant] provides a strong, satisfying conclusion.

Your final workflow is an iterative process:

  1. Paste your script into the text box.
  2. Select a voice you like and the Eleven v3 model.
  3. Set the Stability slider to Natural (around 30-50%).
  4. Add audio tags, punctuation, and emphasis to a paragraph.
  5. Generate the audio and listen.
  6. Refine your tags and regenerate until the delivery sounds perfect.
  7. Download the final MP3 file for use in your video editor.

Conclusion

You've now learned how to bridge the gap between a written script and a polished, professional audio track using AI. This skill is a massive asset for scaling your content production, allowing you to create high-quality marketing materials quickly and consistently without needing to record your own voice for every piece of content.

Key Takeaways:

  • ElevenLabs is an AI tool for generating human-like voiceovers from text.
  • The Eleven v3 model offers the most control over emotion and delivery through audio tags [like this] and punctuation.
  • Your workflow is an iterative cycle of prompting, generating, and refining the audio until it matches your vision.
  • Mastering this tool allows you to create scalable audio assets, a cornerstone of efficient AI-powered marketing.

Preview of the Next Lesson:

While pre-made voices are powerful, establishing a unique and consistent brand voice is even better. In our next lesson, you will clone your own voice using ElevenLabs or PlayHT to create a scalable audio asset for future content. This will enable you to generate audio in your own voice, at scale, for all your future videos, podcasts, and marketing materials.

Can't find a good explanation? Sign up and we'll make it for you

Sign up