Create your own
Lesson illustration

Prioritizing Run Unknowns and Sequencing Work Within Lane Capacity

Welcome back. Your Phase 1 Idea Record now gives the run a bounded problem claim, a real starting contact, a falsifier, credible alternatives, a lane cap, and a Challenger. The next discipline is deciding what to learn first.

A run does not fail because it has unknowns; every new product has them. It fails when the team spends its scarce time resolving low-consequence questions while a single untested condition could make the whole idea untenable. In this lesson, you will turn the assumptions embedded in the Idea Record into a ranked testing order, then fit that order into the lane’s fixed time and money cap.


From a list of worries to a test order

An unknown is a condition that must be true for the current idea to justify its next investment, but for which you lack sufficient, relevant evidence. It is useful to phrase it as a positive, specific claim:

We believe that [specific people] will [observable behaviour or outcome] in [defined context].

For the running example from the previous lesson, “independent consultants miss post-meeting follow-through,” several unknowns are hiding inside the original AI-tool idea:

  • Desirability: Independent consultants with multiple active engagements experience costly follow-through failures often enough to change their current practice.
  • Feasibility: Call material can be turned into accurate, privacy-appropriate action plans with enough human review to be useful.
  • Viability: The value of avoiding missed work supports a price above the dominant operating cost.
  • Distribution: A named channel can reach 100 relevant consultants this month.
  • Usability: A consultant can review, correct, and act on a generated plan without needing training.
  • Ethical or regulatory: The proposed handling of client-call material is permitted and acceptable.

These are not yet facts. At this point, they should be treated as assumptions, even if the Owner has strong experience or an appealing anecdote.

The categories are prompts, not a taxonomy to argue over. Their value is that they prevent the team from generating only customer-demand doubts while overlooking a fatal technical, economic, or governance constraint. Phase 0 has already screened immediate legal, regulatory, and harm issues. But an unresolved compliance path may still be a high-priority feasibility unknown.

The “Map Out Your Hypotheses” matrix places assumptions by their importance and the evidence behind them. The upper-right area contains assumptions that matter greatly but lack evidence, making them candidates for immediate testing.

The matrix is valuable because it separates two questions that teams often merge:

  1. If this is false, how badly does it damage this run?
  2. What direct, traceable evidence do we actually have?

A claim belongs on the “no evidence” side when the support is a hunch, an analogy, a search snippet, an unrelated metric, or a colleague’s confidence. A relevant interview can provide useful directional evidence, particularly for problem discovery, but it is not equivalent to observing repeat use or receiving payment. Evidence must be judged by what happened and how closely it matches the claim, not by how polished the supporting slide looks.

Testing Business Ideas: Assumptions Mapping Webinar

Watch “Testing Business Ideas: Assumptions Mapping Webinar” by David J. Bland. It demonstrates the essential distinction between extracting assumptions, mapping their risk, refining the few that matter most, and selecting a sufficiently cheap next test.

Start with mapping the risks. Focus on the discussion that a high-impact claim with little direct evidence is more urgent than a merely uncertain detail, and on the distinction between interviews and behavioural evidence. Then watch refining hypotheses to see why vague notes must become specific and testable before experiment design. Finish with choosing tests, noting the advice to seek the next increment of evidence cheaply and update the map after learning.


Define “cheaply kill” with four questions

The upper-right quadrant is a good first filter, but the Product Factory adds a harder operational question:

Which uncertainty could produce a decision to STOP with the least additional irreversible spend?

“Cheap” does not mean “quickest activity” in isolation. A two-minute poll is cheap, but it is often too weak to change a serious decision. A useful cheap test is discriminating: a precommitted result would materially change what the team does.

Assess each candidate unknown using these four dimensions.

DimensionWhat to askHigh-priority signal
FatalityIf false, does the current idea stop, simplify radically, or require a new run?The idea cannot reasonably proceed in its current form.
Evidence gapHow much direct, relevant, retrievable evidence exists now?Little or no evidence supports a load-bearing claim.
Test burdenWhat calendar time, money, access, and build effort are required for a decision-grade result?A small test can resolve it before major work begins.
Decision leverageWill a pass or fail change the next decision?Both outcomes have a defined consequence, especially STOP.

A low-priority uncertainty may still be interesting. “Should the first version use a dark or light interface?” could matter later, but it cannot kill the opportunity before a product exists. It should not consume a five-day evidence window.

Conversely, an assumption may be high-impact but not cheaply testable. For example, “will users return unprompted after seven days?” is vital, but requires a usable product and elapsed time. That belongs in Phase 11 validation, not in an attempt to force an early answer through a weak survey.

A practical scoring aid

Use a discussion-based scale, written before choosing tests:

  • Importance : 1 means little consequence; 5 means the current run is untenable if false.
  • Evidence gap : 1 means substantial direct evidence exists; 5 means there is effectively none.
  • Cheapness : 1 means a decision-grade test requires substantial build, money, or external delay; 5 means it can be run quickly with limited spend.

For assumptions that have a genuinely discriminating test, calculate a rough priority score:

This is not a gate and it is not a claim of scientific precision. It is a forcing device for comparing items that otherwise get ranked by whoever speaks most confidently. The numeric gate still comes later: for example, Phase 3’s predefined interview thresholds determine whether the problem passes. The score simply determines which work receives the first hours of the cap.

If two assumptions score similarly, prefer the one that:

  1. Can stop the run sooner.
  2. Has a test that yields stronger evidence.
  3. Must be resolved before another meaningful test can be run.
  4. Has irreversible cost if deferred.

Do not use the score mechanically. A feasibility test with a score of 100 is not automatically “more important” than a core problem test with a score of 75. It may simply mean the feasibility spike is both fatal and inexpensive. The Challenger’s job is to make the trade-off explicit: if we spend the next day on this, what expensive mistake might we avoid?


Build the assumption map before polishing hypotheses

Begin by extracting assumptions rapidly from the Idea Record, the alternatives list, and the Challenger’s initial concern. Capture one claim per note or row. Do not spend an hour perfecting wording for twenty assumptions that will turn out to be low priority.

A short mapping session can be structured as follows:

  1. Set the scope. State the current candidate concept and lane. The question is not whether every imaginable business risk is known; it is which conditions must hold for this run.
  2. Extract assumptions by theme. Spend a short, fixed interval on desirability, viability, feasibility, usability, and remaining ethical or governance constraints.
  3. Map importance and evidence. Discuss each item. Require the person claiming evidence to name the relevant conversation, opened source, data query, or observed behaviour.
  4. Select the top one to three. Choose the high-impact, low-evidence items that can produce a near-term decision.
  5. Refine only those items. Define the specific cohort, observable condition, deadline, evidence source, and failure consequence.
  6. Design the smallest adequate test. Select the least costly test that can actually move the assumption from uncertainty toward a decision.

The key distinction is:

  • Extraction gets worries, beliefs, and hidden premises out of people’s heads.
  • Mapping compares their importance and evidential support.
  • Prioritisation adds the practical cost of resolving them within the lane.
  • Sequencing assigns them a place in the finite calendar and budget.

The Challenger should write the strongest credible “why this fails” assumptions as well as challenge the Owner’s scores. An assumption map created only by the person who proposed the idea will often place their favourite premise too far toward “have evidence.”


Turn the top risks into testable hypotheses

After mapping, sharpen the selected assumptions. A useful hypothesis identifies:

  • Who the claim concerns
  • What must happen
  • In what context
  • How the result will be measured
  • What result ends or changes the run

Here is an illustrative ranking for the consultant follow-through idea. The figures are examples of an ordering method, not evidence or prescribed thresholds.

IDMust-be-true hypothesisCheapest decision-grade test
F-01Call material can produce an accurate, reviewable action plan without unacceptable privacy exposure.554100A time-boxed technical spike using permitted, representative material
D-01At least the required number of independent consultants describe costly follow-through failures before the interviewer names the problem.55375Phase 3 problem interviews with the precommitted threshold
R-01One named channel can expose 100 matching consultants to recruitment or an offer this month.34448A channel inventory plus direct, logged outreach or a named partner’s confirmed access
V-01Avoiding missed follow-up creates sufficient economic value to support the proposed price.45240Quantified current-workaround cost in interviews, followed later by a payment or commitment test
U-01Users can review and act on an output unassisted.45120A prototype or MVP usability test in Phase 11

Two important conclusions follow.

First, the technical spike might rank ahead of full interview analysis because it can kill the AI-based concept within two days. If it fails, there is no reason to spend the rest of the five-day evidence window researching an implementation that cannot meet its core requirement.

Second, D-01 still remains indispensable. A technically feasible output does not establish that anyone has a sufficiently costly problem. The team may begin recruitment while the spike runs, but it must not treat a passing technical result as permission to build.

The Experiment Library compares discovery and validation experiments by cost, setup time, run time, and evidence strength. Use these trade-offs to choose a test that fits the uncertainty and the lane cap rather than defaulting to a preferred research method.

Match the test to the risk and the decision required:

Risk typeUsually appropriate early evidenceCommon mistake
DesirabilityProblem interviews, observation of existing workarounds, behavioural signalsAsking whether someone likes your proposed solution
FeasibilityA bounded engineering spike or technical prototypeWriting a technology assessment instead of making the uncertain thing work
ViabilityCosted economics, pricing context, evidence of spending on alternativesTreating general enthusiasm as willingness to pay
DistributionDirect channel test, reachable-account list, partner confirmationCalling a large theoretical audience “reachability”
UsabilityPrototype or MVP observation when a realistic interaction existsDeclaring a user flow intuitive before anyone uses it

Sequence work inside the lane cap

The lane is a constraint, not a forecast to negotiate later. Its cap determines how many high-priority unknowns can be addressed and how strong the evidence can reasonably become.

For a Lane M run, Phases 3 through 6 share a five-day evidence window. A sound sequence protects the highest-value work first and makes conditional work genuinely conditional.

A Lane M evidence sequence

TimingWorkWhy it happens at this point
Start of Day 1Freeze the priority list, test plans, thresholds, budget allocation, and STOP conditions. Begin recruitment for the required interviews.Prevents thresholds from drifting after evidence arrives; recruitment delay can run while other work starts.
Days 1–2Run the highest-ranked feasibility spike if it is truly capable of killing the concept. Record result and limits.A two-day spike is cheaper than weeks of build or a broad feasibility report.
Days 2–4Conduct the planned customer interviews and record evidence against Phase 3 thresholds.Tests whether the problem exists before solution design or build.
Only if a decision depends on itRun a narrow market, competitive, economics, or reachability check with a written decision trigger.Stops background research expanding to fill the entire cap.
Day 5Calculate actual results, compare them with committed thresholds, score concepts against alternatives, and write the STOP case before the PROCEED case.Converts learning into a gate decision rather than an open-ended research backlog.

This plan permits careful parallelism, not indiscriminate activity. For example, start interview recruitment while an engineer runs a two-day feasibility spike, because recruiting does not commit the team to build. But do not build an MVP, commission market research, and run pricing experiments simultaneously merely because people are available. Work that cannot affect the next gate is outside the current priority.

Stop early only by precommitment

A time-box is not an instruction to use every hour. If a precommitted failure condition is met early, stop the run and cancel remaining work. For instance, a technical spike may show conclusively that the needed data cannot legally be processed in the intended setting.

But avoid stopping prematurely because the first few interviews are discouraging. If the Phase 3 gate requires twelve relevant conversations, the gate result should be calculated from the committed sample unless an early-stop condition was explicitly written or passing has become mathematically impossible. Otherwise, the team is simply replacing one form of motivated reasoning with another.

Keep a lane ledger

For each planned test, record:

Test IDOwnerCalendar allocationCash allocationRequired evidenceSTOP conditionDependency
F-01Engineering ownerUp to 2 daysFixed spike budgetWorking output and risk findingCannot meet agreed quality or data-handling conditionNone
D-01Owner or researcherInterview windowRecruitment incentive budgetInterview records meeting Phase 3 definitionsFewer than the committed thresholdRecruitment started Day 1
R-01OwnerShort bounded checkMinimalTraceable route to 100 ICP membersNo credible named channelOnly if Phase 5 decision needs it

The ledger makes two constraints visible:

  • Calendar cap: work cannot quietly overflow into “one more week.”
  • Money cap: an attractive test cannot consume the whole lane without being deliberately authorised.

For Lane S, apply the same logic at smaller scale: identify the single most decisive uncertainty and test it with the minimum direct contact or operational evidence the one-week cap permits. Do not imitate a Lane M discovery programme for an obvious, reversible improvement.

For Lane L, the same ranking method applies, but legal, safety, data, and irreversible-commitment uncertainties may dominate the map. The cap and review cadence must be stated explicitly, and the required risk register becomes part of the evidence base.


Avoid the two sequencing failures

1. Building before validating

This occurs when “the spike” quietly becomes the product. A valid spike has a narrow question, a hard time limit, defined input material, and a recorded pass or fail result. It does not acquire user accounts, dashboards, integrations, or production architecture unless those are necessary to answer the named technical uncertainty.

A useful check is:

If the spike passes, what do we know? If it fails, what do we stop?

If neither answer changes a decision, it is probably implementation disguised as learning.

2. Researching without shipping

This occurs when every answer generates a new question without a finite decision point. The remedy is not to ban research; it is to define the evidence bar before collection and skip research that cannot affect a gate.

For market and competitive analysis, write a decision trigger such as:

We will analyse direct alternatives only if Phase 3 clears the problem threshold and the Phase 6 concept choice depends on whether a meaningful differentiation exists.

Without that statement, the default is to skip the analysis, note the skip, and preserve time for the most decisive unknown. A broad market report may be informative, but if no possible finding changes the concept decision, it is not currently decision-relevant work.


A reusable prioritisation record

Add this concise section to the Idea Record or Evidence Pack immediately after Phase 1:

UNKNOWN PRIORITY RECORD

Lane cap
Calendar:
Money:
Decision date:

Rank 1
Hypothesis:
Type: Desirability / Feasibility / Viability / Distribution / Usability / Ethical
If false, consequence for this run:
Existing direct evidence and source:
Importance score:
Evidence-gap score:
Cheapness score:
Priority score:
Smallest decision-grade test:
Pass condition:
STOP condition:
Owner:
Time and money allocation:

Rank 2
[Repeat]

Work explicitly deferred
Hypothesis:
Why it is not being tested now:
Phase or gate where it becomes relevant:

Skipped research
Research activity:
Phase 5 or 6 decision it would change:
Reason skipped:

The record should make it possible for someone outside the run to answer four questions quickly:

  1. What could kill this idea?
  2. Why is this unknown ranked above the others?
  3. What evidence would change the decision?
  4. What work was deliberately not done within the cap?

A disciplined run does not try to remove all uncertainty. It identifies the one to three assumptions most able to invalidate the current idea, chooses the smallest adequate tests, and spends the lane budget where a result can change a decision. The map is then updated as evidence arrives; a passing result should move an assumption toward “have evidence,” while a newly discovered risk may take its place at the top.

Next, you will make the highest-priority desirability test operational by defining an initial ideal-customer profile and a recruitment plan for relevant interview participants—while being explicit about what a limited early sample can and cannot support.

Can't find a good explanation? Sign up and we'll make it for you

Sign up