Create your own
Lesson illustration

Launch-Ready A/B Test Plan for Throat Soothing Pops Product Pages

Good to see the implementation brief turned into a measurable experiment. In the previous lesson, you defined a clean product-page change: if the three-pack clears its contribution gate, make it easier to choose through hierarchy, a truthful value comparison, and no new price or discount. This lesson supplies the discipline around that change: who sees it, what counts as success, how long to run it, and what you will do for every outcome.

The objective is not merely to make the three-pack more prominent or to raise AOV. It is to determine whether the new pack presentation increases revenue per eligible product-page visitor without harming conversion or contribution.


1. Treat the page change as one testable decision

Your control is the current Throat Soothing Pops product page. Your treatment is the approved three-pack-emphasis design from the prior lesson:

  • three-pack selected by default;
  • a truthful “Best value per pop” badge, only if its unit price is lower;
  • clear total-pop count, total price, and price per pop for both options;
  • no changed price, discount, shipping threshold, bundle, or upsell logic.

This is a pack-architecture test, not a test of promotions. Keep the rest of the page and all acquisition activity stable. Do not send different paid traffic to each version, change product pricing mid-test, or introduce another product-page release at the same time. Randomly splitting the same eligible visitors keeps CAC comparable between variants.

The required precondition remains in force:

Run the treatment only if the three-pack passed the incremental-contribution gate from Module 1.

If it failed, do not use this test card to justify promoting the three-pack. Keep the neutral presentation while the economics are reconsidered.

A good experiment starts with a mechanism, not just a hoped-for number. Here, the mechanism is straightforward: clear comparison and a sensible default may lead more shoppers to select the three-pack; if that happens without excessive discount exposure or a conversion decline, revenue per visitor can rise.

A Collaborative Template for A/B Tests | by Harlan Harris

Read Harlan Harris’s template to see why a test card needs a stated mechanism, a single decision metric, explicit targeting, and decisions agreed before results appear.

In the “Hypothesis” and “Primary Metric” sections, read the hypothesis framing. Then read “Secondary and Guardrail Metrics,” “Targeting,” “Decision Plan,” and “Allocation.” In “Targeting,” focus on the inclusion decision: define the audience before launch rather than selecting favorable segments afterward.


2. Define the eligible visitor before defining the metrics

The unit of analysis is an eligible visitor, not an order and not a product-page session. This matters because the treatment is shown to visitors before they decide whether to order.

Use this definition:

Eligible visitor: a distinct, non-bot visitor who loads the live Throat Soothing Pops product page during the experiment and is randomly assigned a variant before interacting with the pack selector.

A visitor is counted once, at their first eligible product-page view. Their assignment must remain the same for at least 14 days. If they return tomorrow, they must see the same version. This is often called sticky assignment.

Include and exclude deliberately

IncludeExclude
Mobile and desktop visitors to the live Throat Soothing Pops pageInternal staff, developers, agencies, and QA traffic
New and returning visitorsBots and known test traffic
Visitors from paid, organic, email, direct, and referral sourcesVisitors who cannot receive a stable test assignment because cookies or consent are unavailable, if your tool cannot measure them reliably
Visitors who leave without purchasingPreview URLs, password-protected pages, and staging environments

Do not exclude visitors merely because they chose the two-pack, selected the three-pack, added to cart, or purchased. Those actions are outcomes of the test. Excluding them would make the result biased.

You may later inspect new versus returning visitors, device type, and traffic source as diagnostic segments. But the launch decision must be based on the full eligible population. With a first test, segment results are often too small and noisy to justify separate decisions.


3. Use revenue per eligible visitor as the decision metric

AOV answers, “How much did buyers spend on average?” It ignores everyone who visited but did not buy. A page change that increases AOV while reducing conversion can therefore hurt the business.

Instead, use revenue per eligible visitor, abbreviated here as :

For this test, define net merchandise revenue as:

  • merchandise revenue from the whole attributed order;
  • after automatic discounts;
  • excluding taxes and shipping charges;
  • net of refunds that occur before the analysis is locked.

Count revenue from the entire order, not only Throat Soothing Pops. A shopper may choose a three-pack and add another item; that is part of the commercial effect of the product-page experience.

Use a fixed attribution rule:

Attribute an order when it is placed by a visitor in their assigned browser within 14 days of their first eligible Throat Soothing Pops page view.

If your testing setup cannot reliably connect visitor assignment to Shopify orders for 14 days, use a shorter window that it can measure reliably, such as same-session orders. Write that window into the card before launch and do not alter it later.

The two guardrails

A positive result still is not enough. A guardrail is a metric that can block a launch even when the primary metric rises.

Your first guardrail is conversion rate:

Your second is contribution per eligible visitor, or :

Use the contribution definition already built in Module 1:

Do not subtract automatic discounts twice. They are already reflected when revenue is measured after discounts.

The treatment may increase revenue but decrease contribution if it changes discount qualification, shipping subsidy exposure, payment costs, or the mix of products in the order. catches that outcome.

Track three-pack selection rate and average order value as diagnostic metrics. They explain what happened; they do not decide the launch. For example:

  • Higher three-pack selection but flat may mean conversion fell.
  • Higher AOV but lower may mean the revenue increase was not profitable.
  • No change in selection may mean visitors did not notice or understand the new hierarchy.

How to Conduct A/B Testing in 2026: A Practical Guide ...

Read the sections on metric selection, pre-test analysis, and predetermined decisions. They distinguish a business outcome metric from guardrails and explain why a result should map to an action before the test starts.

In “Selecting the right metrics,” read the discussion from the guardrail definition through its examples. Then, in “How to conduct pre-test analysis,” read the pre-test rationale. Finish with “Mapping your go/no-go decisions for your results,” especially the decision principle.


4. Assign visitors correctly and stop on a planned sample

Use an even 50% control / 50% treatment split. This is the fastest, cleanest allocation for a single, low-risk product-page test.

Your delivery method should meet four requirements:

  1. Randomly assign an eligible visitor when the product page first loads.
  2. Store their assignment in a persistent cookie or testing-platform identifier for 14 days.
  3. Render the same variant at each return visit.
  4. Send an exposure event and connect completed Shopify orders to the assigned variant.

Shopify analytics can help report sales, but it does not itself guarantee valid random assignment. Use your existing Shopify-compatible experimentation tool or a developer-managed feature flag. Do not alternate variants by day, device, traffic channel, or a manually changed theme version. Those approaches mix the design effect with seasonality and traffic-quality differences.

Choose a practical minimum effect

For this first merchandising test, set the minimum practical improvement in at 5% of baseline :

where is historical for eligible Throat Soothing Pops page visitors.

This means you are not trying to “win” because of a trivial uplift that cannot justify keeping a more complex page design. You are testing for an improvement large enough to matter commercially.

Calculate the visitor target

Pull the last 28 days of comparable Throat Soothing Pops page traffic, using the same visitor and revenue definition. Estimate:

  • : baseline revenue per eligible visitor;
  • : the standard deviation of visitor-level revenue;
  • : 5% of .

For a two-variant test using 80% power and a 5% significance level, an approximate visitor target for each variant is:

If your testing platform has a revenue-per-visitor sample-size calculator, use it instead of manually calculating this value, but document the same assumptions: baseline, 5% meaningful lift, 80% power, and 5% significance level.

Your stopping rule is:

Run until each variant has at least eligible visitors and the test has covered at least two complete Monday-to-Sunday weeks. Make the decision once, at that planned endpoint.

The two-week minimum protects against weekday and weekend buying patterns. The sample requirement protects against making a decision from a handful of orders.

Do not look at the result each day and stop whenever the treatment happens to look favorable. You may monitor daily for technical faults, such as a broken selector, incorrect cart variant, price mismatch, or major tracking failure. Those are implementation problems, not performance conclusions.


5. Precommit the decision rules

The decision matrix below captures the central idea: a primary metric needs to improve and guardrails must remain acceptable. Your test card makes the meaning of “pass” explicit rather than relying on judgment after results arrive.

A decision matrix showing that a rollout requires both the primary metric and guardrails to pass; a guardrail failure requires a hold and investigation even if the primary metric is positive.

For this test, use these thresholds:

  • Primary threshold: treatment must improve by at least 5% versus control.
  • Conversion guardrail: treatment conversion rate must not decline by more than 5% relative to control.
  • Contribution guardrail: treatment contribution per eligible visitor must not decline by more than 2% relative to control.

A confidence interval is simply a range of plausible effects given the data. Ask your experiment tool or analyst to report the treatment-minus-control difference and its confidence interval for all three metrics.

Use this action plan:

Result at the planned endpointDecisionAction
The lower end of the 95% confidence interval for improvement exceeds the 5% threshold, and both guardrails meet their limitsLaunchKeep the treatment, then roll it out gradually while monitoring contribution.
Guardrails pass, but the result is uncertain or smaller than the 5% practical thresholdIterateKeep control live. Inspect pack-selection and add-to-cart diagnostics; revise hierarchy or value communication, then test a new version.
The evidence indicates is flat or worse, or either guardrail materially failsStopRoll back to control. Investigate whether the hypothesis, page execution, pricing architecture, or profitability assumptions were wrong.
Selector, cart, tracking, price display, or discount logic is brokenHold immediatelyPause exposure and fix the implementation. Do not interpret commercial metrics from faulty data.

For guardrails, “pass” should be assessed as a non-inferiority check: the data should support that any reduction remains within the agreed tolerance. If your tool cannot calculate this, have an analyst calculate it from visitor-level assigned-variant data. Do not decide from a screenshot of aggregate Shopify orders alone.


6. Copy this launch-ready A/B test card

Complete the bracketed fields before switching on traffic. Most fields are already determined; the remaining fields are your actual page URL, profitability result, historical baseline, and calculated sample size.

Throat Soothing Pops three-pack emphasis: A/B test card

Test name
TSP Pack Architecture v1

Business objective
Grow revenue and order value from Throat Soothing Pops product-page traffic without increasing acquisition spend or reducing contribution.

Pre-launch condition

  • Three-pack incremental contribution versus two-pack: ₹[result from Module 1]
  • Minimum contribution gate: ₹[approved gate]
  • Gate status: Pass required
  • If status is Fail: do not launch this treatment.

Eligible population
Distinct non-bot visitors who first load [live Throat Soothing Pops URL] during the test period, on mobile or desktop.

Exclusions
Staff, agencies, developers, QA traffic, bots, preview/staging pages, and visitors without reliable persistent assignment or order attribution.

Control
Current live Throat Soothing Pops product page: [link to control screenshot]. No three-pack default selection, no new three-pack visual emphasis beyond the current live experience.

Treatment
Approved product-page implementation: [link to treatment screenshot]. Three-pack selected by default, clearly compared with two-pack, and labelled “Best value per pop” only if its lower price per pop is mathematically true. Prices, discounts, shipping thresholds, bundles, and upsells remain unchanged.

Hypothesis
For eligible Throat Soothing Pops product-page visitors, making the profitable three-pack the clear default and showing transparent price-per-pop comparison will increase three-pack selection and increase revenue per eligible visitor by at least 5%, without reducing conversion rate by more than 5% relative or contribution per eligible visitor by more than 2%.

Delivery method
[Testing tool or developer feature-flag owner] randomly assigns each eligible visitor 50% control or 50% treatment on first page render, persists assignment for 14 days, records the exposure, and passes variant identity to attributed Shopify orders.

Traffic allocation

  • Control: 50%
  • Treatment: 50%
  • Assignment: visitor-level and sticky for 14 days
  • No variant-specific acquisition campaign, landing page, or email traffic changes during the test

Primary metric
Revenue per eligible visitor, measured as net merchandise revenue after discounts and before tax and shipping, from whole attributed orders within 14 days, divided by eligible visitors.

Guardrails

  1. Conversion rate: eligible visitors with at least one attributed order divided by eligible visitors. Maximum acceptable treatment decline: 5% relative to control.
  2. Contribution per eligible visitor: total contribution, calculated using the Module 1 cost model, divided by eligible visitors. Maximum acceptable treatment decline: 2% relative to control.

Diagnostic metrics

  • Three-pack selection rate
  • Add-to-cart rate
  • Throat Soothing Pops order mix: two-pack versus three-pack
  • AOV
  • Automatic-discount qualification rate
  • Device and new-versus-returning visitor results, for learning only

Baseline and sample plan

  • Historical period: [dates, previous 28 comparable days]
  • Baseline : ₹[value] per eligible visitor
  • Practical threshold : ₹[0.05 multiplied by ]
  • Historical visitor-level revenue standard deviation : ₹[value]
  • Required eligible visitors per variant : [calculated value]
  • Minimum duration: two complete weeks
  • Planned endpoint: when both the minimum duration and per variant are reached

Stopping rule
Do not make a commercial decision before the planned endpoint. Monitor only technical QA daily. At the endpoint, analyze all eligible visitors according to their originally assigned variant.

Decision owner
[Name: e-commerce owner] with [name: finance or operations reviewer] confirming contribution calculations.

Launch decision
Launch only if exceeds the 5% practical threshold with the required statistical evidence and both conversion and contribution guardrails pass. Roll out in three monitored stages: 25% of eligible traffic for one day, 50% for one day, then 100%, while checking order, contribution, and checkout health.

Iterate decision
If guardrails pass but the result is too uncertain or too small to justify launch, retain control and revise a specific diagnosed issue. Examples include an unclear badge, weak visual hierarchy, or a pack selector that is not visible in the initial mobile view.

Stop decision
If the treatment harms either guardrail beyond its agreed tolerance, or evidence indicates no meaningful improvement, return to control and document the result. Do not try to rescue the result by changing the threshold after seeing the data.


7. The practical launch sequence

Before tomorrow’s implementation work is considered complete, verify these five items:

  1. Economics: the three-pack gate is recorded as Pass using current costs and discount rules.
  2. Experience: control and treatment screenshots are saved for desktop and mobile; the only intended difference is pack architecture.
  3. Measurement: exposure, variant assignment, selected variant, add to cart, completed order, net revenue, and contribution inputs are available.
  4. Sample plan: baseline , practical threshold, required visitors per variant, and planned end date are written into the card.
  5. Decision ownership: one person owns the call; finance or operations confirms the contribution calculation before any rollout.

Key takeaways

A valid A/B test compares randomly assigned, consistently tracked eligible visitors, not whichever orders are easiest to export. For this pack-architecture test, revenue per eligible visitor is the primary metric because it captures both conversion and order value.

Conversion rate and contribution per eligible visitor are guardrails: they can block a rollout even when revenue rises. Set the 50/50 traffic split, sample target, minimum duration, and launch/iterate/stop rules before traffic sees the treatment. That precommitment is what turns a page edit into a business experiment.

This completes the focused sprint: you now have an economics gate, an implementation brief, and an experiment card for testing whether clearer three-pack merchandising can grow AOV without sacrificing profitable growth.

Can't find a good explanation? Sign up and we'll make it for you

Sign up