Does the result hold for different products?
We select one of the first two experiments using criteria fixed in advance and repeat the same procedure with five new products.
- Document version
- Execution plan 01
- Updated
- Design status
- Selection criteria fixed / target experiment not yet selected
- Schedule
- 2026.10.11–10.17
What we want to learn
Research question and hypothesis
Does the content-format difference observed in Weeks 1–2 move in the same direction for new products?
Hypothesis
If the difference is not specific to one product, the primary metric may move in the same direction for new products.
This is an expected direction, not an established finding.Why ask this question?
A single positive result should not become a writing rule without repeating the procedure on new products.
What we will compare
Keep the campaign
change the explanation
Selects one of the first two comparisons and repeats the same procedure on five new products.
Selected baseline condition
Apply the earlier experiment’s baseline material format to new products.
Selected comparison condition
Apply the same experiment’s comparison format to new products.
Place product- and question-level differences from the first experiment beside the new-product results.
Record replications and disagreements by case rather than hiding them in one average.
Controls
Change only the products while keeping question construction, model, no-search setting, repeats, scoring protocol, and cost cap.
Example: travel luggage
Lock either the format comparison or exception-location comparison on October 12, then repeat it with five products not used in the earlier experiment.
How we will proceed
Proposed sequence
Preselect the replication target
Use validity, practical impact, and uncertainty—not only effect size—and lock the reason on October 12.
Preflight new products
Run 24 answers on two new products to verify that material and answer keys follow the same protocol.
Main measurement and comparison
Collect 180 answers for five new products and classify the direction as same, mixed, or opposite to the first experiment.
Do not select the experiment after seeing which effect is largest. Lock the selection reason and version before reviewing results on October 12.
What we will record
Measures and limits
| Measure | Method | Limit |
|---|---|---|
| Primary-metric replication | Calculate accuracy or exception-inclusion rate using the selected experiment’s rule | Do not simply pool the original and new results |
| Product-level consistency | Show direction and error cases for each of five products | Do not declare replication from the average alone |
| Execution match | Check differences in model, questions, scoring, repeats, and cost settings | Changed settings do not count as direct replication |
How we will judge results
Rules fixed in advance
The same direction repeats for new products
Review as a conditional execution-rule candidate
Only some products repeat
Record a conditional result including product and question differences
The direction reverses or disappears
Stop generalizing the first result and include it as a counterexample
Review whether target selection or product replacement favored the result. The same direction alone does not promote the finding to an E3 prescription.
Decisions before launch
Items not yet fixed
- October 12
- Lock replication target A or B and selection reason
- New products
- Five products not used in the prior experiment
- Main measurement
- 24-answer preflight + 180-answer main run
- Cost
- Warning $2 · cap $3
Scope of this plan
- One new-product replication is not external-context replication.
- Because product materials are supplied directly, this is not a web-discovery or citation result.
- A planned experiment remains completed even if the result fails to replicate.
Background and sources
Reviewed guidance
and research rationale
The third comparison in the monthly plan, repeating an earlier design on new products.
Checked 2026-09-26
Read the full backgroundFull explanation and hypothetical example
What we will repeat
We will select either the format comparison from week one or the exception-placement comparison from week two. The choice will be based on experimental validity, practical impact, and uncertainty, and will use five new products.
Measurement scope
The questions, model, scoring rules, and 180-answer scale remain the same. Three repetitions will not be inflated into independent samples; comparisons will be made at the product and question level.
Publication rule
We will record whether the direction is the same, differs by product, or fails to replicate. The result report will include both the selection rationale and the observed differences.