RESEARCH PROPOSALRESEARCH PLAN

Where should a caveat appear so AI does not miss it?

We test whether AI omits conditions more often when exceptions appear at the end of a document rather than beside the relevant claim.

Document version
Execution plan 01
Updated
Design status
Schedule fixed / pre-measurement
Schedule
2026.10.04–10.10

What we want to learn

Research question and hypothesis

Does placing a product restriction beside the related claim reduce missed exceptions in AI answers compared with collecting restrictions at the end?

Hypothesis

Keeping a claim close to its exception may reduce answers that carry over only the benefit while omitting the restriction.

This is an expected direction, not an established finding.

Why ask this question?

If the same facts are reflected differently depending on location, product-description rules can be made more specific.

What we will compare

Keep the campaign
change the explanation

Compares placing the same exception beside its related claim with placing it at the end of the document.

Held constantSame facts · same sentence · same format · same length
A

Exception beside the claim

Place the restriction immediately after the related feature.

B

Exception at the end

Collect the same restriction in cautions at the end of the document.

Measure each condition separately
Primary metricIs the exception stated accurately?

Compare inclusion and accuracy of required restrictions by product and question.

Secondary metricsDoes overall answer quality change?

Also record accuracy, unsupported claims, and cost.

Human response and AI output are measured separately. One does not stand in for the other.

Controls

Keep sentence, facts, document format, questions, model, and repeats constant, changing only exception location.

Example: travel luggage

Compare a version that places “not dishwasher-safe” directly after heat-retention information with one that places the identical sentence in end-of-document cautions.

How we will proceed

Proposed sequence

  1. Lock everything except location

    Use identical sentences for five products and create materials differing only in exception location.

  2. 24-answer preflight

    Verify that storage and scoring hide and compare location conditions correctly.

  3. 180-answer main measurement

    Repeat each condition three times and calculate differences after error review.

Lock the design by October 5; run preflight, main measurement, review, and result writing from October 6 to 10.

What we will record

Measures and limits

MeasureMethodLimit
Exception inclusion rateRecord whether the answer accurately includes answer-key restrictionsExclude exceptions irrelevant to the question from the denominator
AccuracyCompare correct and incorrect fact unitsDo not double-count the exception inclusion rate
Unsupported claimsMark functions or guarantees absent from the materialDistinguish reasonable expression from invented facts
CostSum actual response and scoring-call costsNot the total monthly operating cost

How we will judge results

Rules fixed in advance

Adjacent exceptions reduce omissions across several products

Advance as a new-product replication candidate

Direction differs by product or question

Record by exception type and do not create one rule

Differences remain within repeat variation

Do not name a winner; record that replication is needed

Do not begin main measurement unless the check confirms that location is the only change. Publish completed measurement even if the result is not positive.

Decisions before launch

Items not yet fixed

October 5
Finalize products, exception sentences, and location conditions
Preflight
24 answers · excluded from main analysis
Main measurement
180 answers · compare by product and question
Cost
Warning $2 · cap $3

Scope of this plan

  • This experiment isolates within-document location rather than testing all structure differences at once.
  • Search and indexing are outside scope.
  • Automated scoring is published with the scope of raw-evidence review.

Background and sources

Reviewed guidance
and research rationale

Verisca one-month experiment schedule

The second comparison in the monthly plan, manipulating only the location of exception conditions.

Checked 2026-09-26

Read the full backgroundFull explanation and hypothetical example

What we will compare

The product facts and sentences remain the same. We compare a version that places each exception beside its claim with a version that groups all exceptions at the end of the document.

Measurement scope

The 24-answer preflight and the 180-answer main measurement remain separate. We focus on accuracy and omission of exception conditions while also recording invented benefits and cost.

Publication rule

A completed experiment will be published even if neither condition is better. An experiment that cannot run or cannot be judged will be shown as such, not as complete.