RESEARCH PROPOSALRESEARCH PLAN

Can format change an AI answer when the facts stay the same?

We compare a conventional product description with a question-and-exception format to see whether factual accuracy and omission of conditions change.

Document version
Execution plan 01
Updated
Design status
Schedule fixed / pre-measurement
Schedule
2026.09.27–10.03

What we want to learn

Research question and hypothesis

With equal information, does a format that separates questions and exceptions improve AI answer accuracy and reduce missed conditions compared with ordinary prose?

Hypothesis

Separating questions and exceptions may make relevant facts and restrictions easier to retrieve, improving accuracy.

This is an expected direction, not an established finding.

Why ask this question?

Whether a readable format improves answer accuracy must be tested under matched conditions before it becomes a content-production rule.

What we will compare

Keep the campaign
change the explanation

Compares the same product facts as ordinary prose and as questions with explicit exceptions.

Held constant5 products · 6 questions · same facts and information volume
A

Ordinary prose

Write product facts and restrictions as continuous natural paragraphs.

B

Question-and-exception format

Structure the same facts so customer questions and exceptions are explicit.

Measure each condition separately
Primary metricDoes accuracy differ?

Compare answers with the pre-locked answer key by product and question.

Secondary metricsAre fewer conditions missed?

Record exception omissions, unsupported claims, and response cost separately.

Human response and AI output are measured separately. One does not stand in for the other.

Controls

Hold product facts, questions, model, reasoning setting, no-search condition, and repeats constant, while mixing condition order. Prioritize identical information items rather than character count.

Example: travel luggage

Provide the same facts—such as 500 mL capacity, heat-retention time, and not dishwasher-safe—as prose and as question-and-exception entries.

How we will proceed

Proposed sequence

  1. Lock materials and criteria

    Version five public products, six questions, the answer key, and both formats before viewing results.

  2. Preflight

    Run 24 answers across two products and three questions to verify storage, scoring, information parity, and cost. Exclude them from the main analysis.

  3. Main measurement and review

    Collect 180 answers across three repeats per condition, then compare scoring disagreements and condition samples with raw evidence.

Revalidate the execution path by September 28 and begin main measurement on September 29 only if readiness checks pass.

What we will record

Measures and limits

MeasureMethodLimit
AccuracyCompare fact units in answers with the pre-defined answer keyDo not treat three repeats as independent samples
Condition omissionsRecord whether restrictions and exceptions appear in the answerDo not count information the question did not require
Unsupported claimsMark functions and benefits absent from supplied materialDistinguish wording differences from added facts
CostSum actual response and scoring-call costsExclude agent operations and infrastructure

How we will judge results

Rules fixed in advance

The question-and-exception format moves in the same direction across products and questions

Advance it as a candidate for replication on new products

A difference remains only for some products or questions

Record it conditionally and do not establish a general rule

Storage, scoring, or cost checks fail

Do not begin effect measurement; publish the blocker and recovery deadline

When differences are small or mixed, distinguish no effect from inadequate sample. Results without human review remain an automated-review pilot.

Decisions before launch

Items not yet fixed

September 28
Finalize products, questions, answer key, and live execution path
Preflight
24 answers · excluded from main analysis
Main measurement
180 answers · compare by product and question
Cost
Warning $2 · cap $3

Scope of this plan

  • This is a supplied-material answer experiment, not a test of web discovery or citation.
  • Five public products form a small pilot and do not generalize to every product type.
  • Technical checks using fictional products verify the execution path and are excluded from the result.

Background and sources

Reviewed guidance
and research rationale

Verisca one-month experiment schedule

The first comparison in the monthly research plan running from September 27 to October 26, 2026.

Checked 2026-09-26

Read the full backgroundFull explanation and hypothetical example

What we will compare

We will prepare the same facts for five public products in two formats. One is a conventional product description. The other separates customer questions and exception conditions clearly. Both versions will contain the same information.

Measurement scope

The 24-answer preflight will be reported separately from the main analysis. If the readiness checks pass and expansion is justified, we will collect 180 answers. Accuracy is the primary measure; omitted conditions, unsupported claims, and cost are recorded separately.

Publication rule

We will not publish only favorable results. After checking completion counts, errors, raw evidence, differences, uncertainty, and limitations, we will link the result report to this plan.