Create your own
Lesson illustration

Provenance Tagging and Load-Bearing Assumption Analysis

Welcome back. In the previous lesson, you ranked unknowns by how cheaply they could stop the run and designed bounded tests for the highest-risk assumptions. That work only helps if a later reviewer can tell what supports each claim—and can distinguish a verified observation from a plausible statement that nobody has checked.

This lesson establishes that discipline. You will learn to tag claims in the Idea Record and Evidence Pack with traceable interview, URL, data, or assumption provenance; identify which claims are load-bearing for a gate; and calculate whether unverified assumptions exceed the Product Factory’s 30% ceiling. This is the safeguard that keeps polished reports, AI-generated summaries, and team confidence from impersonating evidence.


Provenance: a claim must have a trail

A claim is a statement in an artifact that could be checked, challenged, or found false. For example:

  • “Eight of twelve relevant interviewees described missed follow-through before being prompted.”
  • “Three interviewees spend at least two hours each week maintaining a workaround.”
  • “A direct outreach list can reach 100 target consultants this month.”
  • “Consultants will pay $49 per month.”

An evidence source is the retrievable material from which a claim was made: an interview recording or transcript, a web page actually opened, or the results of a named query against a named dataset.

Provenance is the connection between the claim and that source. It answers:

Who or what produced this statement, through which activity, based on what underlying material?

The W3C provenance model distinguishes an Agent that performs an Activity from the Entity produced or used by that activity. In a Product Factory run, a participant, researcher, or analyst can be an Agent; an interview or data query is an Activity; and a transcript, query result, or Evidence Pack claim is an Entity.

The diagram is useful because it prevents a common error: treating a report as though it were the evidence itself. An Evidence Pack is an entity derived from evidence. It is not a substitute for the interview record, opened page, or query output that supports it.

Consider this chain:

ElementExample
AgentA target consultant and the researcher
ActivityA recorded problem interview
Evidence entityTranscript int-014
Claim entity“This consultant spends three hours per week reconciling action items.”
ArtifactThe Phase 3 problem section of the Evidence Pack

The tag [int-014] does not prove the claim is representative, important, or sufficient to clear a gate. It does something more basic and essential: it allows a Challenger or decider to retrieve the underlying record and check whether the artifact faithfully represents it.

This distinction matters especially when AI is involved. An AI can summarize int-014, extract candidate claims, and flag inconsistencies across transcripts. But its summary is not fresh evidence. It must remain traceable to the human-collected source.

Product Discovery Meets AI Evals with Teresa Torres

Watch “Product Discovery Meets AI Evals with Teresa Torres” from Hamel Husain. Torres explains why the inputs to an AI product or evaluation cannot simply be invented; they need to be grounded in customer discovery.

Watch grounding eval inputs. Focus on the distinction between generating synthetic material for a technical purpose and guessing the customer personas, use cases, and failure conditions that should shape it. Those underlying customer claims need traceable evidence.


The four provenance tags

The Product Factory uses four permitted tags. Every factual or decision-relevant claim must carry one. If it has no tag, it counts as an assumption.

TagMeaningWhat a reviewer must be able to retrieve
[int-014]A specific interview or conversationThe identifiable interview record, including date, participant fit, notes or transcript, and collector
[url]A page that someone actually openedThe page URL, access date, relevant passage or captured content, and why it bears on the claim
[data: query-name]A named query run against a named datasetQuery definition, dataset identity, query date, output, and relevant inclusion rules
[assumption]Nobody has verified the claim adequatelyThe claim stated honestly as an uncertainty, with a proposed test if it is important

The syntax is intentionally compact. The detail belongs in a source register or evidence repository, not inside every sentence of the Evidence Pack. A tag is a handle that connects the reader to the detailed record.

Use atomic claims

A claim should make one checkable assertion. This prevents a source from appearing to support more than it does.

Too broad

Independent consultants routinely lose client commitments, dislike all current tools, and would pay for an AI follow-up assistant. [int-014]

One conversation is unlikely to establish all three parts, and the sentence merges a problem observation, an alternative assessment, and a payment prediction.

Better

Participant int-014 described missing a promised client follow-up after a meeting. [int-014]

Participant int-014 uses a spreadsheet and weekly review to track commitments. [int-014]

Whether independent consultants would pay for a new workflow remains untested. [assumption]

Atomic claims make disagreements productive. A reviewer can accept the first two statements while challenging the third, rather than rejecting an entire polished paragraph.

A source tag is not a strength rating

A URL can be authoritative, outdated, commercial, irrelevant, or misread. An interview can be direct evidence of one person’s experience but weak evidence of a broad market pattern. A data query can be reproducible but based on a biased dataset.

Provenance answers “where did this come from?” It does not independently answer “is it enough to justify this decision?” The gate, the relevant threshold, and the Challenger’s STOP case answer that second question.

A practical claim register

Use a compact register while drafting the Evidence Pack:

Claim IDExact claimProvenance tagSource locatorLoad-bearing?Limitation
C-01Eight of twelve ICP-matched participants named a follow-through failure before prompting.[data: p3-unprompted-tally-v1]Query against phase3-interview-coding-v1YesEarly sample; not a population estimate
C-02Participant int-014 spends about three hours weekly reconciling action items.[int-014]Transcript and coded noteYesSelf-reported time
C-03A named professional-community partner has offered access to its member newsletter.[int-027]Partner conversation recordYesReach and conversion still uncertain
C-04Consultants will accept automated processing of client-call material.[assumption]NoneYesRequires direct consent and feasibility testing

The blank assumption-log template below is useful as a starting structure because it assigns ownership and validation dates. For this process, extend it with a claim ID, a provenance tag, a source locator, and a load-bearing classification. Otherwise it tracks work without showing what is known.

A blank project assumption log with columns for the assumption, validation owner, due date, update date, and comments. For a Product Factory Evidence Pack, add provenance and load-bearing fields so the log can support a gate review rather than merely list uncertainties.

What makes a claim load-bearing?

Not every sentence in an artifact should affect the assumption calculation. The denominator is limited to load-bearing claims: claims whose truth or evidential status could change a gate decision, concept choice, safety decision, lane, or committed scope.

Use this counterfactual test:

If this claim were false, unsupported, or materially weaker than stated, would the current gate result, decision, or required next action change?

If the answer is yes, it is load-bearing.

Claims that usually are load-bearing

  • The actual counts used to test a numeric gate.
  • Whether interview participants match the stated ICP.
  • Reported workaround cost, time spent, or payment evidence needed for a threshold.
  • The result of the named hardest-unknown spike.
  • A cost, price, or channel-reach claim used in the Phase 5 decision.
  • A legal, privacy, security, or harm claim required to proceed.
  • A factual comparison that changes the Phase 6 concept choice.

Claims that usually are not load-bearing

  • Neutral descriptions that do not affect a decision.
  • Formatting, headings, and methodological instructions.
  • Clearly identified future actions, such as “Run a price test next week.”
  • Contextual detail that does not support the gate result.

“Usually” matters. A statement about a competitor may be non-load-bearing in one run and decisive in another. If your Phase 6 concept selection depends on the claim that no credible alternative provides a particular capability, that competitive claim becomes load-bearing.

The classification must be conservative. A team should not declare a troublesome claim “context” merely to lower the ratio. The Challenger should ask:

Could we still honestly reach the same decision if this statement disappeared?

If not, count it.


Calculate the load-bearing-assumption share

The calculation is straightforward:

For a gate to clear the provenance rule, the result must be at most 30%.

The numerator contains:

  • Every load-bearing claim tagged [assumption]
  • Every load-bearing claim with no tag at all

The denominator contains all load-bearing claims, regardless of whether they are sourced or assumed.

Worked example

Imagine the team is preparing its Evidence Pack for a Lane M opportunity decision. It has identified these eight load-bearing claims:

IDLoad-bearing claimTag
C-01Twelve conversations matched the defined ICP.[data: p3-participant-audit-v1]
C-02Eight of the twelve described the problem unprompted.[data: p3-unprompted-tally-v1]
C-03Three participants spend at least two hours weekly on a workaround.[data: p3-workaround-tally-v1]
C-04At least one participant tried and abandoned a named alternative.[int-019]
C-05Three participants quantified the value of solving the problem.[data: p3-value-tally-v1]
C-06The technical spike met its defined accuracy and data-handling condition.[data: spike-results-v1]
C-07A named channel can reach 100 ICP members this month.[int-027]
C-08The proposed initial price is viable relative to the dominant operating cost.[assumption]

This pack has one load-bearing assumption among eight load-bearing claims:

It meets the provenance ceiling. That does not mean the opportunity automatically clears Phase 5. The team still needs to evaluate whether the remaining assumption is too material for its chosen action and whether all Phase 5 thresholds are met. The ceiling prevents a gate from being supported mostly by unverified premises; it does not convert the sourced claims into good evidence automatically.

Now suppose the team adds this unverified statement:

IDAdded load-bearing claimTag
C-09Target users will accept the required data-processing consent flow.[assumption]
C-10Acquisition through the named channel will cost less than $20 per qualified contact.[assumption]

The revised calculation is:

The artifact remains exactly at the ceiling. If one more load-bearing claim is untagged or assumed, the ratio becomes:

At that point, the artifact cannot clear the gate. The appropriate response is not to relabel assumptions as facts. The team must either collect appropriate evidence within the lane cap, simplify the proposed decision so the claim no longer bears its weight, or stop the run.


Handling interviews, URLs, and data correctly

Interview provenance

An interview tag should identify a real conversation, not a synthesized persona or a paragraph of AI-generated findings. A useful underlying record includes:

  • Interview ID and date
  • How the participant matched the stated ICP
  • Consent and privacy handling appropriate to the context
  • Notes, recording, or transcript
  • The exact observation or quotation supporting the claim
  • Interviewer or collector

Do not write:

“Users consistently struggle with follow-through.” [int-014]

unless int-014 is itself a transparent, named tally derived from a defined set of interview records. One interview cannot support “users consistently.”

Instead, separate the individual record from the aggregate:

Participant int-014 described missing two client commitments in the preceding month. [int-014]

Eight of twelve eligible participants described a follow-through failure before prompting. [data: p3-unprompted-tally-v1]

The first is a participant-level observation. The second is a reproducible aggregate. The query or tally must specify which interview records were included and how “unprompted recognition” was coded.

URL provenance

[url] means the team opened and read the page. A search-result snippet, an AI-generated bibliography entry, or a remembered statistic is not enough.

A strong source record stores:

  • URL and page title
  • Date accessed
  • Relevant passage, table, or archived capture
  • Author or publisher where available
  • Why the material supports this particular claim
  • Important limitations, such as commercial incentives or an unclear methodology

If a page says a competitor offers meeting summaries, it may support:

Competitor A publicly advertises automated meeting summaries. [url]

It does not by itself support:

Competitor A has high customer retention. [assumption]

unless the opened source provides credible evidence for retention.

Data-query provenance

A data tag is appropriate when the claim comes from an identifiable calculation or query. The tag must name the query, while the source register identifies the dataset and preserves the output.

For example:

p3-unprompted-tally-v1 counts eligible Phase 3 interview records in phase3-interview-coding-v1 where unprompted_problem = true.

This is much stronger than a slide that says “67% had the problem.” A reviewer can inspect the sample, inclusion criteria, coding rule, and result.


A short provenance review before every gate

Before the Owner and Challenger write their cases, conduct a bounded review:

  1. Extract claims. Mark factual and decision-relevant statements in the relevant artifact section.
  2. Split merged assertions. Turn a sentence containing several claims into atomic claims.
  3. Attach a permitted tag. Check that each tag resolves to a retrievable interview record, opened page, or named data query.
  4. Classify load-bearing claims. Apply the counterfactual test, with the Challenger reviewing borderline cases.
  5. Count assumptions. Include untagged load-bearing claims in the assumption numerator.
  6. Record the result. Put the numerator, denominator, percentage, and unresolved assumptions in the Decision Log.
  7. Act on the result. If the ratio exceeds 30%, the artifact cannot clear the gate. Decide whether a bounded test, scope reduction, or STOP is justified within the lane cap.

This review should be short because it is performed continuously while evidence is gathered, not reconstructed at the end of a five-day window.

A useful mental rule is:

A source tag lets someone audit a claim. An assumption tag lets someone see the uncertainty. Neither may be hidden.


Key takeaways

Provenance turns an Evidence Pack from persuasive prose into an auditable decision artifact.

  • Tag claims with a specific interview ID, an actually opened URL, a named query against a named dataset, or [assumption].
  • Treat AI summaries, simulated interviews, search snippets, and anonymous “research findings” as unsupported unless they resolve to a retrievable underlying source.
  • Identify load-bearing claims using the counterfactual question: would this claim change the gate or decision if it were false?
  • Calculate the assumption share using only load-bearing claims. Untagged load-bearing claims count as assumptions.
  • An assumption share above 30% prevents the artifact from clearing its gate; the remedy is evidence, simplification, or STOP—not relabeling.

Next, you will initialize the append-only Decision Log and Deviation Log. Those logs preserve not just what was decided, but the threshold, actual figures, STOP case, PROCEED case, decider, and any departure from the standard—making later calibration possible.

Can't find a good explanation? Sign up and we'll make it for you

Sign up