September 3, 2026 · Lance Black
Why your AI visibility score changes month to month

Your AI visibility score rises after you update a pricing page. The tempting conclusion is that the update worked. Then the score falls on the next report, even though nobody touched the site.
Both results are real observations. Neither one, by itself, explains what caused the movement.
AI answers can change because your business changed, competitors and third-party sources changed, or the answer engine changed. A useful report helps you separate those possibilities before you celebrate, panic, or spend another month rewriting pages.
An AI visibility score is a sample
No company can observe every answer produced for every person using ChatGPT, Claude, Gemini, or Perplexity. An AI visibility score summarizes a defined sample of questions, engines, locations, modes, dates, and runs.
That makes the score useful, provided you can inspect what it summarizes. The underlying answers should show whether your business was mentioned, recommended, cited, described accurately, or omitted. They should also show which competitors appeared instead.
Change any part of the sample and the meaning of the comparison can change. Adding branded questions usually makes a known brand appear more often. Switching a model or search mode may change the sources available to the answer. Testing a different location can introduce a different competitive set.
Before interpreting a score change, confirm that the two reports asked comparable questions under comparable conditions. Our guide to measuring whether AI recommends your business explains what a trustworthy baseline should record.
Three things can move the result
A changed answer usually belongs in one or more of these groups.
| What changed | Examples | What to inspect |
|---|---|---|
| Your business's evidence | New product details, clearer service pages, corrected profiles, stronger reviews, earned coverage | Pages and sources connected to the prompts that moved |
| The competitive environment | A competitor launch, new comparison article, updated review page, changed pricing, fresh customer discussion | Competitors and third-party sources appearing in the new answers |
| The engine or measurement | Model update, retrieval change, different mode, prompt edit, location change, ordinary answer variation | Prompt wording, engine, mode, location, date, run count, and complete answers |
The first group is where your work may have helped. The second can move your share of voice even when your own recommendation count stays flat. The third can produce movement that has little to do with either company.
This is why a chart needs annotations. Record when you publish a page, fix crawler access, change pricing, earn coverage, or update an important profile. Without that history, a line moved and everyone gets to invent a reason.
PaperHorse measured the variation directly
In August 2026, PaperHorse ran an internal repeat-run test to see what changed when the question and engine stayed the same.
| Test detail | Sample or result |
|---|---|
| Prompts | 15 |
| Engines | ChatGPT, Claude, Gemini, and Perplexity |
| Repeats per prompt and engine | 10 |
| Total calls | 600 |
| How often a repeated answer cited a different group of websites from the group seen most often | 40% to 85%, depending on the engine |
| Question and engine combinations where the brand was either mentioned every time or never mentioned | 44 of 60 (73%) |
When we repeated the same question ten times, the websites cited often changed. Depending on the engine, 40% to 85% of answers used a different mix of sources from the mix seen most often. Brand mentions were steadier. In 44 of the 60 question and engine combinations, the brand appeared in all ten answers or did not appear in any of them.
This test was deliberately narrow. Fifteen prompts cannot establish a universal volatility rate, and calls through provider APIs may differ from the consumer interfaces customers use. The results still demonstrate a practical risk: a single answer can give a false impression of which sources an engine consistently uses.
Google provides a mechanism for some of this movement. Its documentation says AI Overviews and AI Mode may issue multiple related searches across subtopics and data sources. Google also says those two experiences can use different models and techniques, so their responses and links may vary.
The practical lesson is narrower than "everything is random." The source environment is dynamic, different interfaces can behave differently, and one citation is not a permanent position.
Citation drift is not the same as visibility drift
A cited page supports an answer. A mention names a brand. A recommendation presents that brand as a suitable choice. These events can occur together, but they are not interchangeable.
Your cited URLs could change while your recommendation rate holds steady. Your website could remain a source while the answer recommends a competitor. A new third-party article could mention your brand without linking to your site.
That distinction matters when you diagnose movement. If a citation disappeared, inspect source selection and retrieval. If recommendations fell across a group of buyer questions, inspect the brand claims, competitive differences, and evidence the answers used. If only the combined score changed, open the answers before deciding what the number means.
Decide whether a change deserves action
Use this sequence before assigning a cause or creating work.
1. Confirm the comparison
Check the prompts, engines, modes, geography, dates, run counts, and failed responses. If the measurement changed, label the break rather than presenting one continuous trend.
2. Open the answers that moved
Find the specific prompts where your recommendation status changed. Read what the engines said, which competitors replaced you, and which sources appeared. A ten-point aggregate drop can come from one important topic or several low-value questions. Those situations deserve different responses.
3. Look for persistence
Repeat high-value prompts and look across related questions. One changed answer is a reason to inspect. Movement that persists across repeated runs, several related prompts, or multiple engines is a stronger reason to act.
Cross-engine agreement is especially useful. When the same factual problem appears in ChatGPT, Claude, Gemini, and Perplexity, investigate the shared evidence available to them. When only one engine moves, inspect that engine's answers and source pool before changing the whole site.
4. Check what happened outside the report
Review your release log, content changes, technical fixes, reviews, press coverage, directory profiles, and competitor activity. Search for an observable event that matches the prompts and dates involved.
Sequence alone does not establish causation. If a recommendation appeared after you revised a page, report both facts. Call the page a plausible contributor unless a stronger design isolates its effect.
5. Choose the smallest useful response
Match the action to the evidence. A blocked crawler calls for a technical fix. An incorrect product fact calls for a clear correction on the relevant page and profiles. A recurring competitor on trusted comparison pages may call for outreach or stronger third-party proof. A single inconsistent run may call for no change beyond continued measurement.
The goal is not to react to every wobble. It is to identify the highest-impact gap that the available evidence can support.
Keep before-and-after measurement honest
If you implement a recommendation, record:
- what changed and when;
- which prompts and pages it should affect;
- the evidence supporting the recommendation;
- the expected observable result;
- a stable comparison window;
- other events that could explain movement.
Avoid changing ten things at once if you want to learn which one mattered. Keep an unaffected group of prompts, pages, or locations as a comparison when practical. Preserve every answer so a later review can examine the evidence instead of relying on a scorecard screenshot.
You still may not prove causation. You will have a much better basis for deciding whether to keep the change, expand it, revise it, or wait for more observations.
A monitoring report should lead to a decision
Ongoing monitoring has value because your business, competitors, source landscape, and AI systems keep changing. Collecting more snapshots is not the final outcome. Each report should answer three questions:
- Does AI recommend me? Measure the questions customers ask and keep every answer behind the score.
- Who does it recommend instead? Rank the competitors that appear and show how often they are recommended.
- What should I do next? Turn the observed gaps into a prioritized action plan, with the strongest and most consequential recommendations first.
A changed score tells you where to investigate. The answers, competitors, and supporting evidence tell you what the movement may mean. Prioritization turns that diagnosis into work worth doing.
Get all three questions answered for your business
PaperHorse asks ChatGPT, Claude, Gemini, and Perplexity the questions your customers are asking. It shows where you are winning, where you are missing, and your prioritized next steps to improve. Start free to learn whether AI recommends you, who it recommends instead, and what deserves your attention first.
Start your PaperHorse free trial