Why a Single AI Visibility Score Can Hide Platform Differences
In our comparison of one provider-selection query, ChatGPT and Perplexity surfaced largely different visible source sets, with only one shared domain. Measure visibility on each platform separately first: which queries, which sources, and where your company appears or is absent. A blended score is only meaningful after those platform, query, and source differences are visible on their own.
In this small test, the same provider-selection query produced largely different visible citation sets across the tested Perplexity and ChatGPT experiences. Only one domain appeared in both.
This does not establish a universal preference for either platform. It shows why platform, mode, query, and collection conditions need to remain separate when AI visibility is measured.
What you'll learn#
This article will help you:
- understand why a blended cross-platform visibility score can hide commercially important differences;
- identify which platform, query, source, and collection variables should be reported separately;
- distinguish repeated citation patterns from one-off appearances;
- use visible source differences to guide further investigation without treating them as proof of platform preference or causation.
Key findings#
Finding 1 — The two tested experiences surfaced largely different source sets. For the same query, only one domain (hubspot.com) appeared in both the Perplexity and the ChatGPT visible citations in this test. This is limited to the runs, modes, and date reviewed.
Finding 2 — Each experience showed a repeated core and a more variable tail. A handful of sources appeared in nearly every run; others appeared once or twice. Repeated sampling is what separated the sources that repeated in this sample from the occasional ones.
Finding 3 — The visible source types differed in this sample. Comparison and aggregator pages appeared more often in the Perplexity runs, while product, vendor, and primary pages appeared more often in the ChatGPT thinking-mode runs. These are observed frequencies in this sample, not evidence that either platform favors or rewards a page type.
Finding 4 — Platform name alone does not explain the difference. The two experiences differed in mode, surface, number of runs, and collection conditions, so the observed differences cannot be attributed to "Perplexity vs ChatGPT" as a single variable.
What this means for your company#
A company can appear visible in a blended report while remaining absent from the specific AI experience its buyers use.
Platform-level differences can also change which sources support the answer, which competitors appear, and whether the company's own pages contribute to the result.
The practical decision is not to "optimize for AI" as one channel. It is to identify the buyer questions that matter, measure each relevant platform and surface separately, and then decide whether the opportunity belongs in owned content, customer proof, technical access, or credible third-party presence.
What to measure separately#
For every tracked question, record:
- the exact query and buyer stage;
- platform, surface, and mode;
- date, session, and collection conditions;
- whether the company appears and how it is described;
- whether it is compared or recommended;
- competitors that appear instead;
- visible citation domains and source types;
- which sources repeat across comparable runs.
These fields make the result diagnosable. A single blended score does not.
Detailed observed results#
We ran the same query, "best AEO tools 2026," in two AI search experiences and recorded the visible citations by domain.
What appeared in the Perplexity runs (15 runs):
| Source | How often | Type |
|---|---|---|
| HubSpot | nearly every run | vendor comparison page |
| GetAIRefs | nearly every run | vendor comparison page |
| Vismore | frequently | vendor comparison page |
| AIClicks, Nick Lafferty, Inity | occasionally | mixed source types |
What appeared in the ChatGPT thinking runs (5 runs):
| Source | How often | Type |
|---|---|---|
| Ahrefs, HubSpot, Otterly, Peec, Semrush, Profound, Scrunch | every observed run | mostly tools' own sites |
| Writesonic | nearly every run | tool site |
| AirOps, arXiv | frequently | tool site / primary paper |
| AthenaHQ, Conductor, Evertune, Yext, Brand24, ZipTie, Nightwatch, HiGoodie | occasionally | mixed source types |
Source-type labels are descriptive classifications used for this review. They do not establish independence, authority, or causal influence.
Overlap. Only hubspot.com appeared in both sets. Perplexity's repeated sources (GetAIRefs, Vismore) did not appear in the ChatGPT visible citations, and the ChatGPT repeated sources (Profound, Peec, Otterly, Scrunch, and others) did not appear in Perplexity's.
Breadth per answer. In this sample, the Perplexity answers cited a small, focused set — roughly 3–6 visible sources per answer. The ChatGPT thinking runs cited more broadly — around 11 visible sources per answer, while returning to a core of roughly 8 repeated domains. The difference was not only which domains appeared, but how many.
Core vs. tail. Within each experience, the core sources repeated across most runs; the variation lived in the tail — one-off or low-frequency citations that appeared once or twice. A single prompt run once cannot tell the two apart.
What to investigate next#
- In this Perplexity sample, examine comparison and aggregator coverage further.
- In this ChatGPT mode, examine primary, product, and research-page visibility further.
- Both are hypotheses for the next round of testing, not verified optimization rules.
How we checked this#
We ran the same provider-selection query — "best AEO tools 2026" — in two AI search experiences, each run starting in a fresh session on 2026-06-25. Perplexity was tested in 15 free basic-search runs; ChatGPT was tested in five thinking-mode runs with web search through the API. We recorded every visible cited source by hand and grouped the citations by domain. Because the sample sizes, surfaces, modes, and collection conditions differed, the observed differences cannot be attributed to platform identity alone.
Related Nixal analyses#
What smaller companies can learn from the sources AI cites · Whether you can steer which sources Perplexity uses
FAQ
Why can Perplexity and ChatGPT cite different sources?
The short answer is that citation results are produced by the whole search-and-answer system — not by the model name alone. Perplexity describes its product as a search-first experience that searches the live web, summarizes the retrieved information, and attaches citations to the resulting answer. OpenAI explains that ChatGPT Search may rewrite a user question into one or more targeted searches, review the results, and send additional searches before constructing the response; the ChatGPT setup tested here also combined web search with a thinking model. Differences in retrieval systems, query expansion, source ranking, model mode, reasoning depth, and answer construction can therefore produce different citation sets even when the visible user query is identical. This test cannot assign the observed difference to one specific layer, but the product-level explanation is clear: the two experiences were not running the same search-and-answer pipeline.
What does only one overlapping source tell us?
It tells us that the two tested experiences were drawing on substantially different visible source environments. For a company, that matters because strong citation presence on one platform may not carry into another: a publisher, comparison page, product page, or research asset that appears repeatedly in one experience may never enter the other experience’s visible source set. The practical implication is to report visibility by platform, mode, and buyer question — then examine which sources repeatedly support the answer on each surface. Combining everything into one score would hide where the company is actually present and where it still lacks support.
Should a company optimize for one AI platform?
Usually, no — not before establishing where its buyers actually search and how the company performs on each relevant surface. Most companies need a shared foundation: clear product information, credible evidence, consistent brand facts, technically accessible pages, and independent market support. Platform-specific decisions come after that baseline. In this sample, the Perplexity results suggest investigating comparison and aggregator coverage, while the ChatGPT thinking-mode results suggest investigating primary, product, and research-page visibility — useful starting hypotheses, not permanent platform rules. The investment decision should follow the buyer opportunity and observed source gap, not a generic commitment to one AI platform.
Is this enough data to make a rule?
No. It is enough to form a testable hypothesis. The next step is repeating the same method across more queries, categories, modes, and dates.