Welcome to the first lesson of our module on Designing and Interpreting Experiments.
In the previous module, we focused on evaluating past performance by critiquing marketing reports and dashboards. We learned how to assess them for clarity, relevance, and actionability—essentially, how to make sense of data that has already been collected.
Now, we pivot from looking backward to looking forward. Instead of just analyzing past results, we will learn how to proactively generate trustworthy data through experimentation. As a leader, your ability to oversee a rigorous testing program is what separates reactive teams from innovative, high-growth ones.
This lesson directly addresses the learning outcome: Evaluate A/B test designs for validity and potential biases. We'll explore the foundational principles of a good experiment and create a practical checklist of common threats that can undermine your results. Mastering this will empower you to ask your team the right questions and ensure you're making decisions based on solid evidence, not flawed data.
1. The Core Principles of a Valid Experiment
Before diving into complex statistical concepts, let's start with the fundamentals of experimental design. A/B testing, at its heart, is a controlled experiment designed to answer one question: "Did my change cause an effect?"
To understand what makes a test "good," we need to think about two key concepts: reliability and validity.

To achieve this, every sound experiment is built on two non-negotiable pillars: random assignment and the use of a control group.
Let's watch a short video that explains these core ideas.
Controlled Experiments: Crash Course Statistics #9
This video from CrashCourse explains the fundamental principles that allow a controlled experiment to work. It helps us understand why A/B tests are structured the way they are.
Please watch from 02:10 to 07:02. Focus on these key ideas: Random Assignment (02:10 - 03:54): Why randomly splitting an audience is crucial for creating comparable groups and avoiding selection bias. Control Groups & Placebos (04:47 - 07:02): Why having an untreated 'control' group is essential for isolating the true effect of your change.
As the video explains:
- Random Assignment is the engine of a fair test. By randomly assigning users to see either the original version (Control, or 'A') or your new version (Treatment, or 'B'), you ensure that, on average, the two groups are comparable. This minimizes selection bias, where systematic differences between the groups (e.g., one group has more loyal customers) could skew the results.
- The Control Group is your baseline. It represents the "parallel universe" where you made no change. Without it, you can't know if a change in conversions was due to your new design or something else entirely—like a holiday, a competitor's sale, or just natural variation.
When you evaluate any test design, your first two questions should always be:
- Were the users randomly assigned to each experience?
- Is there a control group to compare against?
If the answer to either is no, it's not a valid A/B test.
2. A Practitioner's Framework for Calling a Winner
While random assignment and control groups are the foundation, they don't guarantee a valid result. Many things can still go wrong during the test. Experienced practitioners use a more holistic checklist before declaring a winner.
This next video provides an excellent, non-technical framework that prioritizes business sense over blind trust in statistical calculators. This is the kind of thinking you want to instill in your team.
A/B Testing & Statistical Significance - 4 Steps to Know How to Call a Winning Test
The video 'A/B Testing & Statistical Significance' from the Testing Theory channel offers a four-step model for confidently calling a winning test. It's a great practical guide for leaders.
Please watch from 01:33 to 09:43. Pay close attention to the four factors the presenter outlines as necessary before declaring a winner. Notice the order in which he presents them.
The video argues that statistical confidence is the last thing you should look at, not the first. A truly valid test result must first satisfy three other criteria:
- Sufficient Data: Have you collected enough conversions in each variation to have a representative sample? A test with 10 conversions versus 2 is not reliable, even if a calculator shows high confidence. The presenter suggests a minimum of 100 conversions per variation, but this number depends heavily on your business.
- Consistent Data: Is the winning variation consistently outperforming the control over time? Early results can be volatile. Look for a stable trend where one variation is clearly and consistently winning for a sustained period (e.g., at least 5-7 days). This helps protect against regression to the mean, where an early fluke result moves back toward the average over time.
- Differentiated Data: Is the lift big enough to be meaningful? If the lift is only 1-2%, it might just be natural "noise" or variance. You need to see a difference that is clearly distinct from the baseline fluctuations of your metrics.
- Statistical Confidence: Only after the first three conditions are met should you look at the statistical numbers (like p-value and confidence intervals). If you have sufficient, consistent, and differentiated data, you will almost certainly have high statistical confidence.
This framework is a powerful tool for you as a leader. When your team presents test results, you can use these four points as your evaluation criteria.
3. A Leader's Checklist of Hidden Biases and Validity Threats
The framework above helps validate a result, but many tests are flawed from the very beginning due to hidden biases in their design. Your role is to spot these threats before the test even runs, or at least to account for them when interpreting results.
The following article from CXL is one of the best resources for identifying these common pitfalls.
How to Minimize A/B Test Validity Threats
This article, 'How to Minimize A/B Test Validity Threats,' provides a detailed, practical list of factors that can secretly sabotage your A/B tests. Think of this as your go-to checklist for critiquing a test design.
Please read the section titled '8 Common Validity Threats Secretly Sabotaging Your A/B Tests.' For each of the 8 threats, focus on understanding what it is and the 'How to manage it...' advice. You don't need to memorize them, but rather become familiar with the types of issues to look for.
Let's group those threats into categories that are easy to remember when you're in a meeting:
1. Technical & Implementation Threats:
- Questions to ask your team: "Did we QA the test setup? Is the revenue tracking firing correctly for all variations? Is there any 'flicker' where users see the old page first?"
- These are basic blocking-and-tackling issues that can completely invalidate a test.
2. Sampling & Timing Threats:
- Questions to ask your team: "Is the test audience representative of our overall user base, or is it biased (e.g., only new users, only paid traffic)? Have we run the test long enough to cover a full business cycle, including weekdays and weekends?"
- These threats compromise External Validity—the ability to generalize your results. A win on a niche audience may not translate to a site-wide rollout.
3. External & Environmental Threats:
- Questions to ask your team: "Are there any other major campaigns (ours or a competitor's) running at the same time? Is there any seasonality or holiday that could be influencing behavior?"
- A test doesn't happen in a vacuum. You must account for the context.
Test your understanding!
One of your social media managers proudly presents the results of a test on a new ad creative.
- Test: They ran the new creative ('B') against the old one ('A') for 48 hours (Thursday-Friday).
- Audience: They targeted a "lookalike" audience on Meta to get fast results.
- Result: Creative B had a 30% higher conversion rate with 98% statistical significance.
They want to immediately shift all ad budget to this new creative. Based on the validity threats we've discussed, what are three critical questions you would ask before approving this?
Show answer
Here are three critical questions to ask, based on the validity threats:
- Selection Bias: "These results are from a lookalike audience. How do we know this creative will perform as well with our core retargeting or existing customer audiences? This audience might be more receptive to novelty."
- Day of Week & Time of Day Effects: "This test only ran for two days at the end of the week. We know user behavior is different on weekends and at the start of the week. We can't be confident this lift will hold over a full 7-day cycle."
- Novelty Effect: "A 30% lift is huge. Is it possible we're seeing a strong initial reaction simply because the creative is new and different? Sometimes performance dips after the novelty wears off. We should monitor this closely if we roll it out."
Your role here is not to dismiss the work, but to inject strategic caution and ensure the decision is robust. A good follow-up would be to suggest a longer test across more representative audiences.
4. Statistical Pitfalls to Watch For
Finally, there are several statistical traps that teams often fall into. You don't need to be a statistician, but you do need to recognize the red flags.
- The "Peeking" Problem: This is the most common mistake. A team launches a test and checks the results daily. The moment it hits 95% significance, they stop the test and declare a winner. This dramatically increases the risk of a false positive. A test must run for its pre-determined duration or sample size, not just until it looks good.
- The Multiple Comparisons Problem: Imagine a test with one control and five new variations. Or a test where the team segments the results by 10 different audiences (new vs. returning, mobile vs. desktop, etc.). The more "shots on goal" you take, the higher the chance of finding a winner purely by luck. When you see results from a many-variation test or heavy post-test segmentation, be extra skeptical. Ask if the team adjusted their significance threshold to account for this (e.g., using a Bonferroni correction, a concept from the provided Adobe resource).
- Confusing Statistical and Business Significance: A test might find a statistically significant lift of 0.5%. While mathematically real, is a change that small worth the engineering effort to implement? We will cover this in more detail in a future lesson.
Conclusion
Evaluating an A/B test design is a core competency for any modern marketing leader. It's your primary defense against making poor decisions based on faulty data. By moving beyond a simple check of statistical significance, you can guide your team toward a more rigorous and intellectually honest experimentation culture.

Key Takeaways:
- A valid test must have random assignment and a control group.
- Use the practitioner's framework to evaluate results: check for sufficient, consistent, and differentiated data before looking at statistical confidence.
- Be a vigilant guardian against hidden threats: question technical setups, biased sampling, improper timing, and external factors.
- Beware of statistical traps like "peeking" and making too many comparisons.
Preview of the Next Lesson:
We've repeatedly mentioned the importance of running a test for a pre-determined duration or sample size. But how do you calculate that? In our next lesson, "Assess the required sample size and duration for an experiment based on statistical power and minimum detectable effect," we will answer that critical question. You'll learn the key inputs needed to plan a test, ensuring you have enough statistical power to detect the changes that matter to your business.