If AI recommends it, will people choose it?
AI recommendations and customer purchases must be measured separately.
- Document version
- Draft 01
- Updated
- Design status
- Design proposal / pre-experiment
- Schedule
- To be determined
What we want to learn
Research question and hypothesis
If AI recommends a product more accurately and persuasively, do actual customers’ understanding and purchase consideration move in the same direction?
Hypothesis
Even when recommendation quality improves, customer appeal ratings and purchase consideration may not change by the same amount.
This is an expected direction, not an established finding.Why ask this question?
AI recommendations, customer beliefs, and purchases are different stages. Separating them shows how far an effect reaches without overstating the next action.
What we will compare
Keep the campaign
change the explanation
Compares changes in AI answers and human choice as separate outcomes.
Current AI answer
Provide the recommendation and explanation produced under the current condition.
Answer using improved material
Use an answer to the same question generated with supplemented product material.
Record factual errors, recommendation reasons, and explanations of poor-fit conditions.
Measure understanding, trust, appeal, and purchase consideration with separate questions.
Controls
Keep the product, price, advertising, question, screen, and recruitment conditions constant, and compare only the answer material.
Example: travel luggage
Show two AI descriptions of a fictional suitcase, then ask separately what customers understood and whether they would include it in their shortlist.
How we will proceed
Proposed sequence
Prepare two answers
Create a current answer and an answer using improved material for the same product and question.
Split presentation
Assign participants without revealing the condition and collect responses with the same questions.
Analyze stages separately
Compare AI quality, customer understanding, and purchase consideration without combining them.
This sequence is a draft. Sample size, repetitions, and decision rules will be fixed before execution.
What we will record
Measures and limits
| Measure | Method | Limit |
|---|---|---|
| AI answer quality | Review errors and omissions against product facts and the recommendation situation | Do not infer customer impact from AI quality alone |
| Customer understanding | Use open responses to check remembered features and suitable users | Do not place the answer inside the question |
| Appeal and purchase consideration | Measure appeal, trust, and shortlist inclusion separately | Not equivalent to an actual purchase |
How we will judge results
Rules fixed in advance
AI quality and customer response both improve
Adopt as a candidate for a follow-up test using behavioral outcomes
Only AI quality improves
Record only an accuracy improvement and make no purchase claim
Customer responses are mixed or lower
Review reading burden and answer expression, then redesign
Participant criteria and a minimum-difference threshold must be set in advance because sample and question design affect customer responses.
Decisions before launch
Items not yet fixed
- Participants
- Target customers and sample size not set
- Answers
- Service, questions, and materials not set
- Decision rule
- Minimum difference and exclusions not set
- Schedule
- Separate candidate from monthly experiments A–C
Scope of this plan
- Intent is not treated as an actual purchase.
- A result from one product or customer group is not directly applied to another market.
- This is a research proposal, not evidence that an experiment has begun or an effect is confirmed.
Background and sources
Reviewed guidance
and research rationale
Product documentation describing recommendations with reasons and tradeoffs based on user preferences and constraints. It is not a study of customer-choice effects.
Checked 2026-09-26
Read the full backgroundFull explanation and hypothetical example
A customer may not choose a product even when AI recommends it
ChatGPT recommended our brand. That is a useful signal, but it does not prove that customers want it more. We do not yet know whether they read or trusted the answer.
Consider travel luggage. An AI may describe a bag as “good for a weekend trip,” while a customer still dislikes the price or design. Someone else may become interested but have no travel plans and therefore make no purchase.
Define the outcome first. If product facts were corrected, measure whether the AI explanation became more accurate. If you want to know whether the campaign’s appeal carried through, read the recommendation rationale. If you want to know whether people want the product more, measure their response.
Recording AI output, customer perception, and purchase separately lets us say what actually improved.
Separate accurate explanation, changed perception, and purchase
| What we want to know | What to check | What this alone cannot tell us |
|---|---|---|
| Is the AI explanation accurate? | Compare it with actual product facts such as size and features | Whether customers like the product |
| Is the recommendation appropriate to the situation? | Check whether the rationale fits the stated context | Whether customers find that rationale persuasive |
| How do customers perceive the brand? | Ask what kind of product it seems to be and what feels different | Whether purchases increased |
| Did customers act? | Record predefined inquiries, comparison choices, or purchases | Whether the campaign caused the change |
A source link and a recommendation are also different. An AI may cite our page but recommend another product, or recommend our product without linking to our page. Record citation and recommendation separately.
Do not put the desired answer in the question
Asking “Recommend our brand for carefree travel” supplies both the brand and the rationale. It becomes difficult to see what the AI would select and why on its own.
Questions should come from real customer inquiries, interviews, or search data. The following are question types, not finalized prompts:
- Ask only about the situation, without naming a brand: Observe which products become candidates for a weekend trip.
- Ask about a specified brand: Observe which trips the bag suits and when it may be inconvenient.
- Compare several brands: Observe how reasons for choosing each bag differ under the same travel conditions.
Named-brand and unprompted questions start from different conditions. Do not combine them into one “recommendation rate.” Creating many questions also does not prove that customers ask them frequently.
Reading supplied material is different from finding it independently
If we attach the old and revised product descriptions to separate prompts, we can compare which one the AI understands more accurately. But because we supplied the material, that test cannot show whether an AI would find the page on the web.
Web discovery requires separate checks: whether the page is public, whether the search service can access it, and whether it is used in an answer. Google also distinguishes eligibility from actual crawling, indexing, and appearance. Meeting the requirements does not guarantee visibility. Google guidance ↗
Record the exact change, retain the original materials and answers, and limit the number of variables changed at once. Hold question wording, AI service, search use, language, and region constant, and save every answer and source.
Also record model changes, price changes, promotions, and new articles during the test. Content changes may not be the only cause of a different answer. The number of repetitions should reflect the difference being tested and the variability of outputs. One correct answer does not establish consistent performance.
Ask actual customers whether they want it
After people review the material, ask what they remember and who they think the product suits, without giving them the company’s preferred language. Then, depending on the objective, ask what feels different, what is attractive, and whether they would include it in a purchase set.
To test whether AI explanations help customers, decide in advance who sees what and in which order. One possible design compares people who see existing material, revised material, and revised material with an AI explanation. The actual number of groups and participants depends on the question and budget.
Asking an AI to “act as a customer” can help find draft problems. It cannot replace results from actual customers.
If only the AI answer improves, check whether customers benefit
| AI explanation and recommendation | Customer response | What to investigate |
|---|---|---|
| Improves in the intended direction | Customers understand and like it more | Whether the result holds for other products and situations |
| Improves in the intended direction | Customers find it less appealing | Whether added explanation weakened the campaign or increased reading burden |
| No confirmed change | Customers understand and like it more | Recognize the customer change while leaving the AI effect unresolved |
| Improvement is unclear | Customer improvement is unclear | Distinguish no effect from insufficient evidence |
Do not call a small difference an improvement without a rule fixed in advance. Define what difference matters and how chance variation will be considered. Without a comparator or predefined rule, report the observed response without claiming that the campaign caused it.
Before testing, write one sentence: “Are we measuring AI accuracy, customer interest, or actual purchase?” A defined outcome makes it easier to decide whether to add material, revise language, or leave the current content alone.
What remains untested
The proposed test has not been run. We do not know the size of any improvement, the participant count required, or the effect on revenue. Before execution, the AI service, questions, participant count, comparator, and stopping rule must be fixed using real customer material.