RESEARCH PROPOSALRESEARCH PLAN

Does the result hold for different products?

We select one of the first two experiments using criteria fixed in advance and repeat the same procedure with five new products.

Document version
Execution plan 01
Updated
Design status
Selection criteria fixed / target experiment not yet selected
Schedule
2026.10.11–10.17

What we want to learn

Research question and hypothesis

Does the content-format difference observed in Weeks 1–2 move in the same direction for new products?

Hypothesis

If the difference is not specific to one product, the primary metric may move in the same direction for new products.

This is an expected direction, not an established finding.

Why ask this question?

A single positive result should not become a writing rule without repeating the procedure on new products.

What we will compare

Keep the campaign
change the explanation

Selects one of the first two comparisons and repeats the same procedure on five new products.

Held constantSame model · question protocol · scoring criteria · 180-answer scale
A

Selected baseline condition

Apply the earlier experiment’s baseline material format to new products.

B

Selected comparison condition

Apply the same experiment’s comparison format to new products.

Measure each condition separately
Replication directionDoes the primary metric move in the same direction?

Place product- and question-level differences from the first experiment beside the new-product results.

Conditional differencesFor which products does the result differ?

Record replications and disagreements by case rather than hiding them in one average.

Human response and AI output are measured separately. One does not stand in for the other.

Controls

Change only the products while keeping question construction, model, no-search setting, repeats, scoring protocol, and cost cap.

Example: travel luggage

Lock either the format comparison or exception-location comparison on October 12, then repeat it with five products not used in the earlier experiment.

How we will proceed

Proposed sequence

  1. Preselect the replication target

    Use validity, practical impact, and uncertainty—not only effect size—and lock the reason on October 12.

  2. Preflight new products

    Run 24 answers on two new products to verify that material and answer keys follow the same protocol.

  3. Main measurement and comparison

    Collect 180 answers for five new products and classify the direction as same, mixed, or opposite to the first experiment.

Do not select the experiment after seeing which effect is largest. Lock the selection reason and version before reviewing results on October 12.

What we will record

Measures and limits

MeasureMethodLimit
Primary-metric replicationCalculate accuracy or exception-inclusion rate using the selected experiment’s ruleDo not simply pool the original and new results
Product-level consistencyShow direction and error cases for each of five productsDo not declare replication from the average alone
Execution matchCheck differences in model, questions, scoring, repeats, and cost settingsChanged settings do not count as direct replication

How we will judge results

Rules fixed in advance

The same direction repeats for new products

Review as a conditional execution-rule candidate

Only some products repeat

Record a conditional result including product and question differences

The direction reverses or disappears

Stop generalizing the first result and include it as a counterexample

Review whether target selection or product replacement favored the result. The same direction alone does not promote the finding to an E3 prescription.

Decisions before launch

Items not yet fixed

October 12
Lock replication target A or B and selection reason
New products
Five products not used in the prior experiment
Main measurement
24-answer preflight + 180-answer main run
Cost
Warning $2 · cap $3

Scope of this plan

  • One new-product replication is not external-context replication.
  • Because product materials are supplied directly, this is not a web-discovery or citation result.
  • A planned experiment remains completed even if the result fails to replicate.

Background and sources

Reviewed guidance
and research rationale

Verisca one-month experiment schedule

The third comparison in the monthly plan, repeating an earlier design on new products.

Checked 2026-09-26

Read the full backgroundFull explanation and hypothetical example

What we will repeat

We will select either the format comparison from week one or the exception-placement comparison from week two. The choice will be based on experimental validity, practical impact, and uncertainty, and will use five new products.

Measurement scope

The questions, model, scoring rules, and 180-answer scale remain the same. Three repetitions will not be inflated into independent samples; comparisons will be made at the product and question level.

Publication rule

We will record whether the direction is the same, differs by product, or fails to replicate. The result report will include both the selection rationale and the observed differences.