Hello! Welcome to your next lesson in the "Designing and Interpreting Experiments" module.
In our previous lesson, we established the strategic importance of balancing exploration (testing for new opportunities) and exploitation (scaling proven winners). You learned that this trade-off is central to managing a marketing portfolio for both short-term efficiency and long-term growth.
Today, we're going to dive into a powerful method that directly addresses this dilemma: Multi-Armed Bandit (MAB) testing. This lesson will show you how algorithms can automate the exploration-exploitation balance, helping you maximize performance in real time. Our goal is for you to be able to explain the principles of multi-armed bandit testing and identify appropriate use cases where it would be more effective than a traditional A/B test.
1. The Core Idea: Automating Exploration and Exploitation
Remember the core trade-off: A/B testing involves a period of pure exploration (splitting traffic evenly to gather data) followed by a period of pure exploitation (sending all traffic to the winner). But what if you could blend these two phases? What if you could start exploiting the likely winner while still exploring the other options, just in case?
This is precisely the problem that multi-armed bandit algorithms are designed to solve. They aim to minimize "regret"—the potential revenue or conversions you lose by sending traffic to an inferior variation during the exploration phase of an A/B test.
To get a clear, intuitive grasp of the concept, let's start with a short video that introduces the multi-armed bandit problem using a simple analogy.
Multi-Armed Bandit : Data Science Concepts
This video from ritvikmath, titled 'Multi-Armed Bandit : Data Science Concepts', provides an excellent introduction to the core problem. It uses a restaurant analogy to define exploration and exploitation and sets up the fundamental challenge that MAB algorithms solve.
Please watch the first part of the video (from 00:00 to 02:41). Focus on understanding the restaurant scenario and how it represents the trade-off between gathering information (exploration) and using the best information you have (exploitation).
As the video explains, the multi-armed bandit problem is about making the best series of decisions over time with incomplete information, constantly balancing the need to learn with the need to earn.
2. How Bandits Differ from A/B Testing
The most critical distinction for you as a marketing leader is understanding how MAB differs from the A/B testing framework you're familiar with.
- A/B Testing: You set a fixed traffic split (e.g., 50/50) and let the test run until you have enough data to declare a statistically significant winner. It's a structured learning phase.
- MAB Testing: The algorithm starts by exploring the variations, but it quickly and automatically begins shifting more traffic towards the variation that is performing better. It learns and optimizes at the same time.
This image perfectly illustrates the difference in traffic allocation over time:
To solidify this distinction, let's turn to a practical guide from CXL.
Guide to Multi-Armed Bandit: When to Do Bandit Tests
The article 'Guide to Multi-Armed Bandit' offers a clear, marketing-focused comparison between the two methods.
Please read the section titled 'The practical differences between A/B testing and bandit testing'. Focus on the 'Explore-exploit' subsection and the diagrams that visualize how each approach handles exploration and exploitation.
The key takeaway is that A/B testing is designed to answer the question, "Which of these variations is better, and by how much?" MAB, on the other hand, is designed to simply deliver the best possible outcome over the duration of the activity by minimizing traffic to underperforming options.
3. A Simple Bandit Strategy: Epsilon-Greedy
You don't need to code these algorithms, but understanding the logic of a simple one will demystify how they work. The most common approach is called Epsilon-Greedy.
The logic is straightforward:
- Set a value for epsilon (ε), typically small, like 0.1 (or 10%). This is your "exploration budget."
- For 10% of your users (the "epsilon"), the algorithm will explore by showing them a randomly chosen variation.
- For the other 90% of users, the algorithm will exploit by showing them the variation that has the highest conversion rate so far.
Let's return to the video to see this in action.
Multi-Armed Bandit : Data Science Concepts
The same video now explains the Epsilon-Greedy strategy and compares it to more naive approaches.
Please continue watching from 02:41 to 09:20. This part first shows why 'explore only' or 'exploit only' are bad strategies. Then, it explains how the Epsilon-Greedy method provides a much better balance and results in lower 'regret'.
The power of this simple rule is that it guarantees you never stop learning entirely (thanks to the epsilon), but you heavily favor the option that's proven to be best. This is the mechanism that allows the algorithm to "earn while it learns."

4. Strategic Use Cases: When to Use a Bandit
Now for the most important part of this lesson: When should you, as a leader, choose a bandit test over a traditional A/B test? The decision depends on your goal.
Use a Multi-Armed Bandit when your primary goal is to maximize performance and minimize opportunity cost during the activity.
This makes it ideal for two main categories of marketing activities:
1. Short-Term Campaigns and Promotions
- Use Case: Headlines for a news story, subject lines for a time-sensitive email, or creative for a week-long Black Friday campaign.
- Why it Fits: By the time a traditional A/B test concludes, the campaign might be over! A bandit test allows you to automatically shift traffic to the winning ad creative or headline in real-time, maximizing clicks or sales when it matters most. You're not trying to learn a timeless principle; you're trying to win right now.
2. Automated, Continuous Optimization ("Set it and forget it")
- Use Case: Choosing which ad creative to show from a pool of 10 options in an evergreen campaign, or deciding which recommended products to display on a homepage.
- Why it Fits: For high-volume, low-risk decisions that need to be made constantly, a bandit can act as an automated optimization agent. It continuously learns and adapts, ensuring that your best-performing assets get the most exposure without requiring constant manual analysis.
The following resource provides an excellent breakdown of these use cases.
Multi-Armed Bandit vs AB Testing—A Guide
The Braze article 'Multi-Armed Bandit vs AB Testing—A Guide' offers a concise, practical checklist for deciding when to use MAB.
Please read the section 'When to use multi-armed bandit optimization'. Pay close attention to the 'Gut check (yes/no)' list. This gives you a fantastic mental model for quickly assessing if a situation is right for a bandit test.
Test your understanding!
Your team is planning the annual holiday promotion, which runs for the first three weeks of December. They have designed four different hero banners for the homepage, each with a unique offer and design. The Head of Ecommerce wants to run a standard A/B/C/D test to see which one performs best. What is your recommendation and why?
Show answer
You should recommend a Multi-Armed Bandit test instead of a traditional A/B/n test.
Your reasoning would be:
- Time-Sensitivity: The campaign is short and high-stakes. With a traditional A/B/n test, you would spend a significant portion of the three weeks sending 75% of your valuable holiday traffic to what will likely be underperforming variations.
- Goal is Maximization, Not Just Learning: The primary objective during the holiday season is to maximize revenue, not to achieve perfect statistical certainty about the exact lift of each banner.
- Minimize Opportunity Cost: A bandit algorithm will quickly identify the early top-performer(s) and start sending more traffic their way, maximizing conversions from day one. It minimizes the "regret" of showing a weak offer during a peak sales period.
You can frame it as: "Let's use a bandit to let the data automatically optimize our homepage for revenue throughout the campaign. After the holidays, we can analyze the results to learn which offers resonated, but during the campaign, our priority should be performance."
5. Advanced Application: Contextual Bandits for Personalization
The MAB concept can be taken a step further. A standard bandit finds the best single option for everyone. A Contextual Bandit goes deeper: it finds the best option based on the context of the user.
Context could include:
- The user's device (mobile vs. desktop)
- Their location
- The time of day
- Whether they are a new or returning visitor
This video explains the concept of contextual bandits and how they enable a basic form of personalization.
Contextual Bandits : Data Science Concepts
This follow-up video, 'Contextual Bandits : Data Science Concepts', explains how adding 'context' makes bandit algorithms much more powerful and relevant for personalization.
Please watch this entire video (around 10 minutes). It builds directly on the previous one, showing how a 'one-size-fits-all' approach can be improved by considering user attributes. Focus on the strategic implication: moving from optimizing for the average user to optimizing for different user segments.
As a leader, this is a powerful concept to have in your toolkit. When your analytics team presents results, you can ask the critical next-level question: "This is the best-performing creative on average, but is it the best for every customer segment? Could we use a contextual bandit to personalize the experience for our mobile users versus our desktop users?"
Conclusion: Adding Bandits to Your Strategic Toolkit
Multi-armed bandit testing is not a replacement for A/B testing; it is a complementary tool for a different job. Understanding when to use each is a hallmark of a sophisticated marketing leader.
Key Takeaways:
- MAB automates the exploration-exploitation trade-off, allowing you to "earn while you learn."
- The primary benefit of MAB is minimizing opportunity cost (regret) by dynamically shifting traffic to winning variations.
- Use MAB for short-term campaigns (headlines, promos) or continuous, automated optimization where maximizing immediate performance is the goal.
- Use traditional A/B testing when you need high statistical certainty to validate a strategic hypothesis and understand the "why" behind a lift before making a major change (e.g., a site redesign, a pricing change).
- Contextual Bandits are a powerful extension that allows you to move from one-size-fits-all optimization toward real-time personalization.
Preview of the Next Lesson:
Whether you're planning a series of A/B tests or setting up a bandit, you'll always have more ideas than you have resources to test them. How do you decide which ideas are worth pursuing? In our next lesson, we will get practical and apply a prioritization framework (e.g., ICE, RICE) to an experimentation backlog, ensuring your team focuses its efforts on the tests with the highest potential business impact.