RESEARCH PROPOSALRESEARCH PLAN

Where is content selected for an AI answer?

Good writing alone is not enough. A page must pass each stage—access, discovery, retrieval, evidence selection, and answer use.

Document version
Execution plan 03
Updated
Design status
Waiting at indexing gate / citation measurement not started
Schedule
Resume after the indexing gate passes

What we want to learn

Research question and hypothesis

For documents on the same topic and of similar length, does including sourced numbers increase citation in actual AI answers?

Hypothesis

When both treatment and control documents are discoverable, the document with verifiable numbers may be cited more often.

This is an expected direction, not an established finding.

Why ask this question?

Prior work mainly studied environments where candidate sources were already supplied. A new site requires separate validation after crawl, indexing, and retrieval.

What we will compare

Keep the campaign
change the explanation

Compares a document containing verifiable numbers with a qualitative version carrying the same meaning.

Held constantSame topic · same structure · similar length · same discovery path
A

Document with numbers

Include verified statistics and numbers in the body.

B

Qualitative document

Describe the same facts qualitatively without numbers.

Measure each condition separately
Prerequisite gateWere both documents discovered?

Count a pair as valid only when the treatment and control URLs are each found.

Citation measurementWhich document is used more often?

Run three repeats per service condition only when at least three valid pairs exist.

Human response and AI output are measured separately. One does not stand in for the other.

Controls

Match title, description, H1, structure, crawl setting, and inbound-link wording across four pairs, changing only whether the body contains numbers.

Example: travel luggage

Explain the same topic with the same structure and similar length; one version uses sourced numbers and the other expresses the same meaning qualitatively.

How we will proceed

Proposed sequence

  1. Import index state

    Import Search Console and search-discovery observations while preserving source and check time for each URL.

  2. Determine valid pairs

    Check whether at least three of four pairs have both URLs independently discovered.

  3. Collect citations by condition

    Only after the gate passes, separately repeat Google AI Mode, AI Overview, and OpenAI Responses API collection.

Gate collection is stopped because the required index-import structure is not deployed in the production database. The last completed observation is not used as evidence of no effect or confirmed non-indexing.

What we will record

Measures and limits

MeasureMethodLimit
Discovered URLsDistinguish URL/title search from Search Console evidenceNot found in search is not proof of internal non-indexing
Valid pairsCount only pairs where both treatment and control were foundBelow three pairs, withhold the effect decision
Citation ratePreserve cited URLs by question version, service condition, and repeatDo not combine service conditions
Use of numeric claimsCheck whether numbers were actually used as answer evidence when citedSeparate link presence from content use

How we will judge results

Rules fixed in advance

Fewer than three valid pairs

Do not start citation measurement; keep the result undecidable

Numeric documents outperform beyond repeat variation among valid pairs

Record as an observation for a new-domain condition and plan replication

Results are mixed or both conditions receive zero citations

Separate no effect from an upstream discovery problem

Do not convert sitemap processing or discovered-page totals into URL-level indexing. Do not calculate a zero observation as a 0% citation rate.

Decisions before launch

Items not yet fixed

Current blocker
Production DB index-import structure not deployed
Resume condition
Completed gate run and at least 3 valid pairs
Last observation
0/8 found · 0/4 valid pairs · includes partial errors
Result state
Citation measurement not started · no effect decision

Scope of this plan

  • The last observation is one search-discovery snapshot, not a definitive indexing state.
  • A new-domain result cannot be generalized directly to established large sites.
  • High similarity between experiment documents may cause one canonical to be consolidated.

Background and sources

Reviewed guidance
and research rationale

GEO: Generative Engine Optimization

Prior work comparing numbers, sources, and other features in an environment with candidate sources supplied. Verisca Lab tests replication with real discovery and indexing stages.

Checked 2026-09-26

Read the full backgroundFull explanation and hypothetical example

Start with the sequence

AI search services do not share one internal architecture, and their exact ranking signals are not public. Still, public documentation and research suggest a practical path that can be observed:

Query expansion → access → discovery and indexing → candidate retrieval → source selection → answer use → brand mention and recommendation → visit and action

When an early stage fails, later improvements cannot be evaluated. A citation does not prove that a source substantially shaped the answer, and appearing in an answer does not prove recommendation or a visit. Each stage therefore needs its own measure.

1. One question expands into several searches

Google says AI Overviews and AI Mode may use “query fan-out” to expand one question into searches across subtopics and multiple data sources.

What to do: Make the topic, audience, and conditions of a document explicit, and cover the follow-up questions people actually ask.

How it may work: When the meaning of a query and a document is close, the document may have more opportunities to enter relevant retrieval sets. Public evidence does not establish that question-style headings, keyword-style headings, or one specific outline is universally superior.

2. Crawlers can access the body

Google states that a page must be indexed and eligible for snippets to become a supporting link in its AI features. OpenAI advises publishers not to block `OAI-SearchBot` if they want content included in ChatGPT Search summaries and snippets. Perplexity similarly recommends allowing `PerplexityBot` for search visibility.

What to do: Do not block search crawlers in `robots.txt`, firewalls, or CDNs, and place important information in static text.

How it may work: A system must retrieve the body before it can use the page as a search candidate or source. Crawl access is a prerequisite, not a proven ranking or citation bonus.

3. The document is discovered and indexed

Internal links and sitemaps are routes for discovering new URLs. Google recommends making content easy to find through internal links, while also stating that satisfying technical requirements does not guarantee crawling, indexing, or visibility.

What to do: Avoid orphan pages and check internal links, canonical URLs, sitemaps, and index status separately.

How it may work: A document outside a search index cannot become a candidate for a search-grounded answer. Submitting a sitemap alone should not be interpreted as a ranking or citation increase.

4. It is selected as a source for the question

Search systems do not place every indexed document in an answer. They choose candidates using relevance, search quality, freshness, source characteristics, and other signals whose exact weights vary and are not public.

What to do: Use titles, introductions, and body text that directly match the scope of the question. Prefer original research, experience, or data over summaries of other pages.

How it may work: Google says its search-based generative features use its core ranking and quality systems to retrieve relevant, fresh pages, and it suggests that original, non-commodity content is more likely to matter over time. This is official guidance, not a published causal estimate for any single factor.

5. Evidence suitable for an answer is selected

The KDD 2024 GEO study, in an environment where candidate sources were already supplied, found that statistics, citations, relevant quotations, and fluency could increase a source’s share and position in generated answers. Keyword repetition showed no benefit in the same experiment.

What to do: Replace vague language with verifiable numbers, link primary sources, and attribute quotations clearly. Test numbers, quotations, and sources separately rather than changing all of them at once.

How it may work: Numbers, quotations, and sources provide concrete evidence units a model can reuse. The study did not test the full process of discovering and indexing a new page on the open web, so replication in live services remains necessary.

6. Evidence is absorbed into answer sentences

The 2026 study “From Citation Selection to Citation Absorption” measures source-link selection separately from the extent to which source content shapes an answer’s sentences, evidence, and structure. Longer, clearly structured, semantically relevant pages with extractable definitions, numbers, comparisons, and procedures were associated with greater answer influence.

What to do: Separate definitions, comparison criteria, procedures, denominators, and conditions so readers can distinguish them. Use clear headings and readable prose without copying a format mechanically.

How it may work: An answer generator can find and reorganize the needed evidence units more easily. The study reports an association from large-scale observation; it does not establish that short paragraphs or a particular HTML structure caused the outcome.

FeatGEO also reported that optimizing document-level semantic and structural features outperformed changing a few isolated words in its benchmark. It too requires replication before being generalized as a rule for every live service.

7. Brand choice and visits are separate outcomes

A source link does not prove that a brand was described favorably or recommended. Even a recommendation does not prove that a person visited the source or bought anything.

What to do: Record URL citation, content absorption, brand mention, recommendation direction, source visit, and subsequent action as separate outcomes.

How it may work: Stage-level measurement helps distinguish content problems from brand-perception problems and visit-motivation problems. Observations suggest that third-party brand mentions can affect AI answers, but their causal effect and required volume are not established.

What is relatively well supported

  • Crawling and indexing are prerequisites for entering search-grounded answer candidates. Access alone does not guarantee citation.
  • There is no evidence that a special AI-only file or schema is required for Google’s AI features. Google states that no additional technical requirements apply beyond its existing search fundamentals.
  • When candidate sources are already provided, statistics, citations, relevant quotations, and fluent writing have improved source visibility in generated answers.
  • Keyword repetition did not show a stable advantage in that experiment.
  • The presence of a citation link and a source’s actual influence on answer content are different measures.

What remains unknown

  • The exact source-selection signals and weights used by ChatGPT Search, Google AI Mode, AI Overviews, and Perplexity
  • Whether earlier findings hold for Korean-language content, the Korean market, and new domains
  • The independent effects of question headings, top summaries, FAQs, tables, short paragraphs, and schema
  • Whether a recent date itself matters, or whether actual content updates drive selection
  • The causal effect and required repetition of third-party brand mentions on brand selection in answers
  • The conditions under which answer exposure and citation lead to visits, inquiries, or purchases

What the Lab will test

The execution order follows the points where failure blocks later stages, not the popularity of a tactic.

  • Access gate: Allowing or blocking search crawlers. Crawlability and citation are measured separately.
  • Discovery gate: Internal links, sitemaps, and other discovery routes. Time to first indexing and failure states are recorded.
  • Candidate retrieval: The effect of semantically aligned titles, subheadings, and body text on retrieval. A question-style title remains a separate variable.
  • Evidence selection: Numbers, source links, direct quotations, and original research are changed one at a time.
  • Answer absorption: Definition-, comparison-, and procedure-led structures are compared with continuous prose; fluency is isolated as a separate variable.
  • Freshness: Actual content updates and date changes are separated for time-sensitive questions.
  • Brand outcomes: First-party URL citation, third-party mention, recommendation, and visit are measured separately.

The current numerical Lab experiment is still checking the indexing gate. Citation rates will not be calculated until enough paired documents are indexed. An undiscovered page is not recorded as evidence of no effect.

Evidence and limitations

The official documents describe each company’s own systems but do not disclose complete ranking formulas. The papers use different datasets, dates, engines, and metrics. We do not average their results, and we do not treat them as implementation rules until the Lab reproduces them under comparable conditions.