Fractional CTO
Back to Research

Research

AI Engines Don't Agree on Sources

Seven months of ChatGPT, Claude, Gemini and Perplexity answers to 119 buyer questions in one B2B services niche

September 29, 2026 · 15 min read

I track what AI engines cite for the questions buyers ask in my own niche, fractional CTO services. This report covers seven months (February 27 to September 26, 2026), four engines, 119 buyer questions in English and Spanish, and 24,742 source links after cleaning.

Ask the four engines the same question in the same week, and ChatGPT shares 4–6% of its cited domains with any of the other three. Those three share 14–18% with each other. Each engine repeats 29–55% of its own sources from one week to the next, so this isn't noise: the engines are drawing from different parts of the web.

Two limits apply to everything below. This is one niche, the one I sell into, so my own sites are in the data (4–8% of each engine's sources); every finding was rerun without them and none changes direction. And these are the sources each engine's API returns, which isn't the same as what the consumer apps show people. The methodology has the details.

24,742Source links
2,942Answers with sources
1,749Domains cited
7Months tracked

1. ChatGPT cites a different web from the other three

For every question two engines answered in the same weekly run, I compared the sets of domains they cited. The measure is Jaccard similarity: shared domains divided by all domains either one cited. A score of 1 means identical lists; 0 means nothing in common. The same-engine rows compare one engine with itself across weeks, on the same question.

Comparison (July 27–September 26)Domain overlapQuestion-weeks compared
ChatGPT vs Claude4.4%256
ChatGPT vs Perplexity5.1%442
ChatGPT vs Gemini5.7%123
Claude vs Perplexity13.6%295
Claude vs Gemini16.8%131
Gemini vs Perplexity17.6%140
Same engine, week to week29–55%20–80 questions

ChatGPT gives short source lists (3.7 domains per answer, against 9–16 for the others), and short lists overlap less by construction. So I also checked the overlap coefficient, which divides by the shorter list. ChatGPT still comes last: 13–32% against 31–42% for the other pairs. The gap holds without my own domains, and over the full seven months (ChatGPT vs Claude: 4.8% across 794 question-weeks).

The favorites differ too. These are Claude's most distinctive sources, counted over the 794 question-weeks both engines answered:

DomainClaude answers citing itChatGPT answers citing it
gofractional.com22210
cto.academy1145
arc.dev970
aikenhouse.com801
fractionalctos.org711

I'm not showing the mirror table for ChatGPT, because the domain at the top of it is my own Spanish site, ctofraccional.com (134 answers against Claude's 0).

Why it matters. "AI visibility" as one number averages four different things. A page that Claude, Gemini and Perplexity all cite can still be invisible to ChatGPT, and the reverse is more likely still. Track each engine on its own.

2. Claude and Gemini often don't search at all

The collector asks every engine every question. Whether the answer comes back with sources is up to the model, and the engines differ by a factor of four:

EngineAnswers with sourcesEnglishSpanish
Perplexity100%100%100%
ChatGPT94.1%92.4%96.7%
Claude56.0%61.1%48.4%
Gemini26.6%28.3%23.9%

July 27 to September 26: 527 question-runs per engine (314 English, 213 Spanish).

Gemini is the extreme case. It searched on 6.8% of informational questions ("what does a fractional CTO do?"), against 38.6% of commercial ones. For most definitional questions it answers from what the model already knows, and no page gets cited.

Two caveats. ChatGPT's API ran with an instruction to always search and cite, and the others had none, so part of ChatGPT's 94% is my prompt. And Claude searches more in English than in Spanish (61% against 48%), which matters if you publish in both.

Why it matters. Before you optimize a page to be cited for a question, check whether the engine searches for that kind of question at all. For Gemini and definitional questions, mostly it doesn't.

3. Reddit is an engine choice, not a category trait

Reddit is the source AEO advice talks about most. In this niche, whether it gets cited depends almost entirely on the engine:

SourceAnswers citing RedditQuestions with Reddit
ChatGPT (API)0 of 1,3070 of 110
Claude (API)0 of 9270 of 78
Gemini (API)5.5% of 1816 of 31
Perplexity (API)23.3% of 52744 of 96
Google top 10, 13 head terms80.9% of 2,420 daily snapshots13 of 13

The same questions that get Reddit from Perplexity get none from ChatGPT or Claude, and Google puts Reddit in the top 10 for almost every head term in this niche (between 1% and 11% of listings on each). The Reddit threads exist. Two of the engines don't use them. Otterly's Claude citation study found Claude citing Reddit 0 times across categories, which fits.

One thing I can't explain yet. On the 39 questions Perplexity answered in every run, the number of answers citing Reddit went from 19 on August 8 to 1 on September 26, falling in six straight runs. Google's Reddit share didn't move over the same weeks. The counts are small, so read it as an observation, not a trend. It looks like the mid-August drop Promptwatch reported for ChatGPT Search (from 3.83% of citations to 0.52%), but my ChatGPT data can't show that one: the API never cited Reddit to begin with.

Why it matters. "Should I invest in Reddit for AI visibility?" depends on which engine you care about more than on your category. For Perplexity it can pay off; for the ChatGPT and Claude APIs it did nothing in seven months.

4. ChatGPT links to homepages; the others link to articles

I sorted every cited URL into one of four types. Each URL gets exactly one: a homepage first, then a service or landing page, then an article, then anything else.

EngineHomepageService / landingArticleOtherURLs
ChatGPT27.9%27.0%38.3%6.9%4,974
Claude6.3%18.0%66.1%9.6%8,588
Perplexity3.8%27.4%59.0%9.7%9,530
Gemini1.0%32.2%58.4%8.5%497

Gemini covers three September runs only; its older links had expired before I could resolve them.

28% of ChatGPT's links go to homepages (28.3% without my own sites). ChatGPT and Perplexity answered the same questions here, so the gap isn't a difference in question mix.

Where do ChatGPT's homepage links come from? Mostly from answers of the form "here are some firms you could hire", which link each firm's homepage. More and more of those answers also carry a Google Maps link for each firm: 0% of ChatGPT answers in February, 14.9% in August.

Why it matters. For Claude, Gemini and Perplexity, the page that gets cited is an article that answers the question. For ChatGPT, being a firm it names matters as much as the article, and the link it gives people is your homepage. Make sure that page says clearly what you do.

5. Each question has a sticky winner and a shuffling tail

For each question an engine answered at least four times, I looked at two things: how often its most-cited domain showed up, and how much of the whole list repeated from week to week.

EngineTop domain presentList repeats week to weekQuestions
Perplexity99% of answers37.7%80
Claude96%55.0%39
Gemini93%28.6%20
ChatGPT82%32.5%62

The top slot barely moves. Everything below it reshuffles every week. For context, Google's top 10 for this niche's head terms repeats about 52% week to week, so ChatGPT, Gemini and Perplexity shuffle more than Google and Claude about as much. (Different queries on each side, so that's context, not a head-to-head test.) This fits SparkToro's finding that AI recommendations are inconsistent, with one qualification: the head is stable.

Why it matters. Getting cited once is easy and means little. A single week's check will show you in the tail one week and gone the next. The slot worth tracking is the one that holds for weeks.

Smaller things worth knowing

  • Every ChatGPT API link is tagged. All 7,973 ChatGPT URLs carry utm_source=openai, every month. Claude and Perplexity tag none. These are API responses, not clicks, so they say nothing direct about what shows up in your analytics.
  • A pinned API model doesn't see consumer releases. Across the consumer GPT-5.5 release (May 22–23), the ChatGPT API's sources changed no more than in ordinary weeks. That's expected for a pinned gpt-4o, and it's a limit of every API-based tracker, this one included.
  • Citation shares drift. For "Best fractional CTO in Latin America", three sources held 32% of citations through early May. Over the full seven months and four engines it's 20.6%, with a different top three.

What this means if you're trying to be cited

  • Measure each engine separately. With 4–6% overlap between ChatGPT and the rest, a blended score tells you very little.
  • For ChatGPT, be a firm it can name, and have a homepage that says what you do. For the other three, publish specific articles that answer the question.
  • Check whether the engine searches for your kind of question before you build a page for it. Gemini mostly doesn't for definitional questions, and Claude skips search about half the time.
  • Decide on Reddit by engine. Perplexity uses it; the ChatGPT and Claude APIs didn't once in seven months.
  • Judge citation wins over weeks. The top slot is stable; the rest of the list is a lottery.

Methodology

  • Window: February 27 to September 26, 2026, in 45 runs. My earlier report covered the first 2.5 months with ChatGPT and Claude only. Daily until mid-March, weekly after that. No runs on March 16–21, August 1–2 (a schedule change) or September 5–6 (a crash).
  • Engines, all through their developer APIs, one fresh single-turn request per question, no location sent:
    • ChatGPT: gpt-4o (an undated alias) with web search, and the instruction "Always search the web before answering. Cite your sources."
    • Claude: claude-sonnet-4-5-20250929 with web search limited to one search per answer, no system prompt.
    • Gemini: gemini-flash-latest (a moving alias) with Google Search grounding, from July 25.
    • Perplexity: sonar, from July 27.
  • A "source" differs by engine. ChatGPT: links cited in the answer. Claude: every result of its one search, whether or not the answer used it. Gemini: grounding sources, one per domain. Perplexity: its source list. So this study measures what each API returns as sources.
  • Questions: 119 prompts (84 English, 35 Spanish) about hiring a fractional CTO, which is what I sell. 91 were generated by an LLM from my own notes. Questions that went four checks without citing me or a tracked competitor were retired (30 of them), which tilts the set toward questions where I show up.
  • My own sites (fractionalcto.com.ar, ctofraccional.com) are 8.3% of ChatGPT's source domains, 4.3% of Claude's, 8.0% of Gemini's and 3.8% of Perplexity's. Every finding was rerun without them.
  • Unit: a domain appearance, meaning one domain cited in one answer. It evens out engines that list several URLs from one site. Cross-engine comparisons use only questions both engines answered in the same run, from July 27 onward.
  • Cleaning: repeated links within an answer removed (Perplexity stored every source twice). Google Maps place links and search-engine pages excluded (821 rows). Gemini's redirect links resolved where they still worked (497 of 1,653).
  • Answers without sources aren't stored, so I rebuilt the denominators from the runs. The logs show 62 failed Claude calls and none for the other engines; the rest of the missing answers are the model choosing not to search.
  • Google data: the daily top 10 for 13 head terms in this niche, US English (Spanish terms moved to Google Argentina on September 8). It's context for the topic, not a question-by-question comparison.

The full dataset is available as a CSV: 37,346 raw rows with engine, model, question, language, cited and resolved URL, domain, URL type, whether the domain is mine, and why a row was excluded. Drop the excluded and duplicate rows to get the cleaned 24,742.

Open questions

  • How far do the consumer apps differ from their APIs? Everything here is API output.
  • Does Perplexity's Reddit decline continue, and does anything replace it?
  • What does Google show for the same long questions? That needs a question-by-question SERP pull, which I don't have yet.

About this report

This data comes from an SEO/AEO monitoring agent I built to track my own visibility across Google and AI engines. I publish it because the questions buyers ask AI engines about professional services get little public data.

Geographic engagement map showing page sessions by city across the Americas, Europe, Asia, and Oceania
The same agent tracks geographic engagement, content gaps, keyword clusters, and SERP movement. The AEO module is one slice of a broader monitoring system.

Questions, corrections, or requests for additional cuts: connect on LinkedIn.

Ezequiel Actis Grosso

Ezequiel Actis Grosso

Fractional CTO

I help startups turn architecture, technical risk, and engineering constraints into business decisions. 25+ years shipping software.

Follow on LinkedIn

Newsletter

Want the next report?

Subscribe to get notified when I publish the next update and other research.

Next step

Want to talk about how this applies to your site?

I'm a fractional CTO and I built the agent that produced this data. Happy to share the methodology or run a custom cut for your domain.

Book a Free 30-min Call