Can format change an AI answer when the facts stay the same?
We compare a conventional product description with a question-and-exception format to see whether factual accuracy and omission of conditions change.
- Document version
- Execution plan 01
- Updated
- Design status
- Schedule fixed / pre-measurement
- Schedule
- 2026.09.27–10.03
What we want to learn
Research question and hypothesis
With equal information, does a format that separates questions and exceptions improve AI answer accuracy and reduce missed conditions compared with ordinary prose?
Hypothesis
Separating questions and exceptions may make relevant facts and restrictions easier to retrieve, improving accuracy.
This is an expected direction, not an established finding.Why ask this question?
Whether a readable format improves answer accuracy must be tested under matched conditions before it becomes a content-production rule.
What we will compare
Keep the campaign
change the explanation
Compares the same product facts as ordinary prose and as questions with explicit exceptions.
Ordinary prose
Write product facts and restrictions as continuous natural paragraphs.
Question-and-exception format
Structure the same facts so customer questions and exceptions are explicit.
Compare answers with the pre-locked answer key by product and question.
Record exception omissions, unsupported claims, and response cost separately.
Controls
Hold product facts, questions, model, reasoning setting, no-search condition, and repeats constant, while mixing condition order. Prioritize identical information items rather than character count.
Example: travel luggage
Provide the same facts—such as 500 mL capacity, heat-retention time, and not dishwasher-safe—as prose and as question-and-exception entries.
How we will proceed
Proposed sequence
Lock materials and criteria
Version five public products, six questions, the answer key, and both formats before viewing results.
Preflight
Run 24 answers across two products and three questions to verify storage, scoring, information parity, and cost. Exclude them from the main analysis.
Main measurement and review
Collect 180 answers across three repeats per condition, then compare scoring disagreements and condition samples with raw evidence.
Revalidate the execution path by September 28 and begin main measurement on September 29 only if readiness checks pass.
What we will record
Measures and limits
| Measure | Method | Limit |
|---|---|---|
| Accuracy | Compare fact units in answers with the pre-defined answer key | Do not treat three repeats as independent samples |
| Condition omissions | Record whether restrictions and exceptions appear in the answer | Do not count information the question did not require |
| Unsupported claims | Mark functions and benefits absent from supplied material | Distinguish wording differences from added facts |
| Cost | Sum actual response and scoring-call costs | Exclude agent operations and infrastructure |
How we will judge results
Rules fixed in advance
The question-and-exception format moves in the same direction across products and questions
Advance it as a candidate for replication on new products
A difference remains only for some products or questions
Record it conditionally and do not establish a general rule
Storage, scoring, or cost checks fail
Do not begin effect measurement; publish the blocker and recovery deadline
When differences are small or mixed, distinguish no effect from inadequate sample. Results without human review remain an automated-review pilot.
Decisions before launch
Items not yet fixed
- September 28
- Finalize products, questions, answer key, and live execution path
- Preflight
- 24 answers · excluded from main analysis
- Main measurement
- 180 answers · compare by product and question
- Cost
- Warning $2 · cap $3
Scope of this plan
- This is a supplied-material answer experiment, not a test of web discovery or citation.
- Five public products form a small pilot and do not generalize to every product type.
- Technical checks using fictional products verify the execution path and are excluded from the result.
Background and sources
Reviewed guidance
and research rationale
The first comparison in the monthly research plan running from September 27 to October 26, 2026.
Checked 2026-09-26
Read the full backgroundFull explanation and hypothetical example
What we will compare
We will prepare the same facts for five public products in two formats. One is a conventional product description. The other separates customer questions and exception conditions clearly. Both versions will contain the same information.
Measurement scope
The 24-answer preflight will be reported separately from the main analysis. If the readiness checks pass and expansion is justified, we will collect 180 answers. Accuracy is the primary measure; omitted conditions, unsupported claims, and cost are recorded separately.
Publication rule
We will not publish only favorable results. After checking completion counts, errors, raw evidence, differences, uncertainty, and limitations, we will link the result report to this plan.