The query is not one fixed keyword
These prompts can produce different candidate sets:
Best carry-on backpack
Best carry-on backpack under $150
Best carry-on backpack for a 16-inch laptop
Best carry-on backpack available in Calgary this week
A store does not occupy one position across all four.
The environment changes
Answers can vary with:
- platform and model
- search or shopping mode
- user location and language
- conversation history
- product price and inventory
- source freshness
- merchant eligibility
- provider experiments
Retrieval is the process of selecting sources or catalog records before an answer is generated. Different retrieval systems create different inputs.
A score can still be useful if it is defined
A test set can report the proportion of completed prompts with a mention or citation. That is a measurement of that test set during that period.
For example:
Canonical store domain cited in 6 of 18 completed prompts,
using 20 fixed prompts across two platforms in Canada on July 27.
Two provider failures were excluded.
This is defensible because the denominator and scope are visible.
It should not become:
The store ranks 33% in AI.
Measure observations, not a fictional index
Use stable prompt IDs, timestamps, source URLs, market, completion status, mentions, citations, competitors, and fact errors. Preserve raw evidence.
When a merchant changes a product page, compare the same prompt set before and after. The result may suggest an effect; it does not prove causation without broader controls.
The detailed method is in Sampling AI visibility without inventing a ranking.
Community discussion
Add to the article
Ask a technical question, share a storefront result, or challenge a conclusion with evidence.