Why asking ChatGPT about yourself gives you a false reading
The instinct is to open ChatGPT and type your own brand name. That test is close to worthless, and it fails in three separate ways at once.
First, naming your brand in the prompt guarantees it appears in the answer — you have asked a leading question. What matters commercially is whether you surface for a prompt that never mentions you, like "best analytics tool for a Shopify store."
Second, if you are logged in, ChatGPT's memory and any custom instructions are in play, and you have probably discussed your own company with it before. You are measuring your own conversation history, not the model.
Third, model output is non-deterministic. The same prompt run twice can produce different brands in different orders. Any method that reads a single response as a verdict is measuring sampling noise.
The five-step manual method
This is the process to run before you automate anything. It takes about two hours the first time and gives you a baseline number.
- 1
Write 20-30 prompts a real buyer would type
Draft prompts across the funnel without naming your brand: category discovery ("best X for Y"), comparison ("X vs Y for small teams"), problem-first ("how do I stop Z from happening"), and constraint-led ("cheapest X that integrates with Shopify"). Twenty is the floor for a stable rate.
- Pull real language from your sales call notes and support tickets.
- Include the prompts your competitors would rank for, not just yours.
- Avoid brand names entirely in this set — keep those for a separate reputation set.
- 2
Run each prompt in a clean, logged-out session
Open a private window and run each prompt without signing in. If you must be signed in, disable memory and clear custom instructions first. Never reuse a thread between prompts — earlier turns bias later answers.
- One prompt per fresh thread.
- Note the model version; GPT-4o and newer reasoning models answer differently.
- Location affects results — note the country you ran from.
- 3
Log four things per response, not one
For each answer record: whether your brand appeared at all, its ordinal position in the list, the sentiment of the sentence describing it, and every URL cited. The citation list is the most actionable column in the sheet.
- Position matters — being fifth in a list of five is close to invisible.
- Log competitor names in the same row so you get share of answer for free.
- 4
Repeat the full set across at least three separate days
Run the same prompt set on day one, day three, and day seven. Three passes is the minimum needed to separate a genuine absence from a single unlucky sample. Average the mention rate across passes.
- 5
Convert the log into one number and re-baseline monthly
Divide the prompts where you appeared by the total prompts run to get your mention rate. That percentage — not any individual answer — is the metric. Re-run the identical set monthly so the number stays comparable.
The columns your tracking sheet needs
If you are doing this in a spreadsheet, these are the non-negotiable columns.
Prompt text
Verbatim, so the run is reproducible next month.
Date and model version
Answers shift materially between model releases.
Brand mentioned (Y/N)
The raw input to your mention rate.
Position in list
First mention versus fifth is a different commercial outcome.
Sentiment of the mention
Being named as the expensive option is not a win.
Cited URLs
The domains the model leaned on. This is your content roadmap.
Competitors named
Turns the same sheet into a share-of-answer report.
Where the manual method runs out
| Manual spreadsheet | EvidentlyAEO | |
|---|---|---|
| Prompts per run | 20-30 before fatigue sets in | Hundreds, scheduled |
| Engines covered | ChatGPT only, realistically | ChatGPT, Gemini, Perplexity, Claude, Copilot |
| Run cadence | Whenever someone remembers | Daily, automatic |
| Sampling noise | Handled by hand, if at all | Averaged across repeated runs |
| Citation attribution | Copy-paste URLs | Sources ranked by influence on answers |
| Competitor share of answer | Manual tally | Computed per prompt and per topic |
Run the manual method first regardless. It teaches you what your prompt set should contain, which is the part no tool can do for you.
What to do with the citation column
The mention column tells you where you stand. The citation column tells you what to do about it.
Group every cited URL by domain and count frequency. You will usually find that a small number of third-party sources — a review aggregator, one or two industry publications, a comparison roundup, a Reddit thread — account for most of the citations across your prompt set. Those domains are the ones shaping the answer.
Where a competitor appears and you do not, open the cited page and check whether you are listed on it at all. A large share of AI invisibility is not a model problem; it is an absence from the handful of pages the model trusts for that category.
Separate owned pages from earned pages
Citations pointing at your own domain mean your content is already legible to the model — extend that pattern to the topics where you are missing.
Citations pointing at third parties are an outreach and PR task, not a content task. Treat them as two different workstreams with different owners.