Where should a caveat appear so AI does not miss it?
We test whether AI omits conditions more often when exceptions appear at the end of a document rather than beside the relevant claim.
- Document version
- Execution plan 01
- Updated
- Design status
- Schedule fixed / pre-measurement
- Schedule
- 2026.10.04–10.10
What we want to learn
Research question and hypothesis
Does placing a product restriction beside the related claim reduce missed exceptions in AI answers compared with collecting restrictions at the end?
Hypothesis
Keeping a claim close to its exception may reduce answers that carry over only the benefit while omitting the restriction.
This is an expected direction, not an established finding.Why ask this question?
If the same facts are reflected differently depending on location, product-description rules can be made more specific.
What we will compare
Keep the campaign
change the explanation
Compares placing the same exception beside its related claim with placing it at the end of the document.
Exception beside the claim
Place the restriction immediately after the related feature.
Exception at the end
Collect the same restriction in cautions at the end of the document.
Compare inclusion and accuracy of required restrictions by product and question.
Also record accuracy, unsupported claims, and cost.
Controls
Keep sentence, facts, document format, questions, model, and repeats constant, changing only exception location.
Example: travel luggage
Compare a version that places “not dishwasher-safe” directly after heat-retention information with one that places the identical sentence in end-of-document cautions.
How we will proceed
Proposed sequence
Lock everything except location
Use identical sentences for five products and create materials differing only in exception location.
24-answer preflight
Verify that storage and scoring hide and compare location conditions correctly.
180-answer main measurement
Repeat each condition three times and calculate differences after error review.
Lock the design by October 5; run preflight, main measurement, review, and result writing from October 6 to 10.
What we will record
Measures and limits
| Measure | Method | Limit |
|---|---|---|
| Exception inclusion rate | Record whether the answer accurately includes answer-key restrictions | Exclude exceptions irrelevant to the question from the denominator |
| Accuracy | Compare correct and incorrect fact units | Do not double-count the exception inclusion rate |
| Unsupported claims | Mark functions or guarantees absent from the material | Distinguish reasonable expression from invented facts |
| Cost | Sum actual response and scoring-call costs | Not the total monthly operating cost |
How we will judge results
Rules fixed in advance
Adjacent exceptions reduce omissions across several products
Advance as a new-product replication candidate
Direction differs by product or question
Record by exception type and do not create one rule
Differences remain within repeat variation
Do not name a winner; record that replication is needed
Do not begin main measurement unless the check confirms that location is the only change. Publish completed measurement even if the result is not positive.
Decisions before launch
Items not yet fixed
- October 5
- Finalize products, exception sentences, and location conditions
- Preflight
- 24 answers · excluded from main analysis
- Main measurement
- 180 answers · compare by product and question
- Cost
- Warning $2 · cap $3
Scope of this plan
- This experiment isolates within-document location rather than testing all structure differences at once.
- Search and indexing are outside scope.
- Automated scoring is published with the scope of raw-evidence review.
Background and sources
Reviewed guidance
and research rationale
The second comparison in the monthly plan, manipulating only the location of exception conditions.
Checked 2026-09-26
Read the full backgroundFull explanation and hypothetical example
What we will compare
The product facts and sentences remain the same. We compare a version that places each exception beside its claim with a version that groups all exceptions at the end of the document.
Measurement scope
The 24-answer preflight and the 180-answer main measurement remain separate. We focus on accuracy and omission of exception conditions while also recording invented benefits and cost.
Publication rule
A completed experiment will be published even if neither condition is better. An experiment that cannot run or cannot be judged will be shown as such, not as complete.