RESEARCH PROPOSALRESEARCH PLAN

If AI recommends it, will people choose it?

AI recommendations and customer purchases must be measured separately.

Document version
Draft 01
Updated
Design status
Design proposal / pre-experiment
Schedule
To be determined

What we want to learn

Research question and hypothesis

If AI recommends a product more accurately and persuasively, do actual customers’ understanding and purchase consideration move in the same direction?

Hypothesis

Even when recommendation quality improves, customer appeal ratings and purchase consideration may not change by the same amount.

This is an expected direction, not an established finding.

Why ask this question?

AI recommendations, customer beliefs, and purchases are different stages. Separating them shows how far an effect reaches without overstating the next action.

What we will compare

Keep the campaign
change the explanation

Compares changes in AI answers and human choice as separate outcomes.

Held constantSame product · same question · same presentation environment
A

Current AI answer

Provide the recommendation and explanation produced under the current condition.

B

Answer using improved material

Use an answer to the same question generated with supplemented product material.

Measure each condition separately
AI answerIs it accurate and appropriate to the situation?

Record factual errors, recommendation reasons, and explanations of poor-fit conditions.

Customer responseDo understanding and choice change?

Measure understanding, trust, appeal, and purchase consideration with separate questions.

Human response and AI output are measured separately. One does not stand in for the other.

Controls

Keep the product, price, advertising, question, screen, and recruitment conditions constant, and compare only the answer material.

Example: travel luggage

Show two AI descriptions of a fictional suitcase, then ask separately what customers understood and whether they would include it in their shortlist.

How we will proceed

Proposed sequence

  1. Prepare two answers

    Create a current answer and an answer using improved material for the same product and question.

  2. Split presentation

    Assign participants without revealing the condition and collect responses with the same questions.

  3. Analyze stages separately

    Compare AI quality, customer understanding, and purchase consideration without combining them.

This sequence is a draft. Sample size, repetitions, and decision rules will be fixed before execution.

What we will record

Measures and limits

MeasureMethodLimit
AI answer qualityReview errors and omissions against product facts and the recommendation situationDo not infer customer impact from AI quality alone
Customer understandingUse open responses to check remembered features and suitable usersDo not place the answer inside the question
Appeal and purchase considerationMeasure appeal, trust, and shortlist inclusion separatelyNot equivalent to an actual purchase

How we will judge results

Rules fixed in advance

AI quality and customer response both improve

Adopt as a candidate for a follow-up test using behavioral outcomes

Only AI quality improves

Record only an accuracy improvement and make no purchase claim

Customer responses are mixed or lower

Review reading burden and answer expression, then redesign

Participant criteria and a minimum-difference threshold must be set in advance because sample and question design affect customer responses.

Decisions before launch

Items not yet fixed

Participants
Target customers and sample size not set
Answers
Service, questions, and materials not set
Decision rule
Minimum difference and exclusions not set
Schedule
Separate candidate from monthly experiments A–C

Scope of this plan

  • Intent is not treated as an actual purchase.
  • A result from one product or customer group is not directly applied to another market.
  • This is a research proposal, not evidence that an experiment has begun or an effect is confirmed.

Background and sources

Reviewed guidance
and research rationale

OpenAI guide to shopping research

Product documentation describing recommendations with reasons and tradeoffs based on user preferences and constraints. It is not a study of customer-choice effects.

Checked 2026-09-26

Read the full backgroundFull explanation and hypothetical example

A customer may not choose a product even when AI recommends it

ChatGPT recommended our brand. That is a useful signal, but it does not prove that customers want it more. We do not yet know whether they read or trusted the answer.

Consider travel luggage. An AI may describe a bag as “good for a weekend trip,” while a customer still dislikes the price or design. Someone else may become interested but have no travel plans and therefore make no purchase.

Define the outcome first. If product facts were corrected, measure whether the AI explanation became more accurate. If you want to know whether the campaign’s appeal carried through, read the recommendation rationale. If you want to know whether people want the product more, measure their response.

Recording AI output, customer perception, and purchase separately lets us say what actually improved.

Separate accurate explanation, changed perception, and purchase

What we want to knowWhat to checkWhat this alone cannot tell us
Is the AI explanation accurate?Compare it with actual product facts such as size and featuresWhether customers like the product
Is the recommendation appropriate to the situation?Check whether the rationale fits the stated contextWhether customers find that rationale persuasive
How do customers perceive the brand?Ask what kind of product it seems to be and what feels differentWhether purchases increased
Did customers act?Record predefined inquiries, comparison choices, or purchasesWhether the campaign caused the change

A source link and a recommendation are also different. An AI may cite our page but recommend another product, or recommend our product without linking to our page. Record citation and recommendation separately.

Do not put the desired answer in the question

Asking “Recommend our brand for carefree travel” supplies both the brand and the rationale. It becomes difficult to see what the AI would select and why on its own.

Questions should come from real customer inquiries, interviews, or search data. The following are question types, not finalized prompts:

  • Ask only about the situation, without naming a brand: Observe which products become candidates for a weekend trip.
  • Ask about a specified brand: Observe which trips the bag suits and when it may be inconvenient.
  • Compare several brands: Observe how reasons for choosing each bag differ under the same travel conditions.

Named-brand and unprompted questions start from different conditions. Do not combine them into one “recommendation rate.” Creating many questions also does not prove that customers ask them frequently.

Reading supplied material is different from finding it independently

If we attach the old and revised product descriptions to separate prompts, we can compare which one the AI understands more accurately. But because we supplied the material, that test cannot show whether an AI would find the page on the web.

Web discovery requires separate checks: whether the page is public, whether the search service can access it, and whether it is used in an answer. Google also distinguishes eligibility from actual crawling, indexing, and appearance. Meeting the requirements does not guarantee visibility. Google guidance ↗

Record the exact change, retain the original materials and answers, and limit the number of variables changed at once. Hold question wording, AI service, search use, language, and region constant, and save every answer and source.

Also record model changes, price changes, promotions, and new articles during the test. Content changes may not be the only cause of a different answer. The number of repetitions should reflect the difference being tested and the variability of outputs. One correct answer does not establish consistent performance.

Ask actual customers whether they want it

After people review the material, ask what they remember and who they think the product suits, without giving them the company’s preferred language. Then, depending on the objective, ask what feels different, what is attractive, and whether they would include it in a purchase set.

To test whether AI explanations help customers, decide in advance who sees what and in which order. One possible design compares people who see existing material, revised material, and revised material with an AI explanation. The actual number of groups and participants depends on the question and budget.

Asking an AI to “act as a customer” can help find draft problems. It cannot replace results from actual customers.

If only the AI answer improves, check whether customers benefit

AI explanation and recommendationCustomer responseWhat to investigate
Improves in the intended directionCustomers understand and like it moreWhether the result holds for other products and situations
Improves in the intended directionCustomers find it less appealingWhether added explanation weakened the campaign or increased reading burden
No confirmed changeCustomers understand and like it moreRecognize the customer change while leaving the AI effect unresolved
Improvement is unclearCustomer improvement is unclearDistinguish no effect from insufficient evidence

Do not call a small difference an improvement without a rule fixed in advance. Define what difference matters and how chance variation will be considered. Without a comparator or predefined rule, report the observed response without claiming that the campaign caused it.

Before testing, write one sentence: “Are we measuring AI accuracy, customer interest, or actual purchase?” A defined outcome makes it easier to decide whether to add material, revise language, or leave the current content alone.

What remains untested

The proposed test has not been run. We do not know the size of any improvement, the participant count required, or the effect on revenue. Before execution, the AI service, questions, participant count, comparator, and stopping rule must be fixed using real customer material.