CrediGeo research · July 21, 2026
State of AI Software Recommendations 2026
How much of AI’s B2B-software advice is built from the vendors’ own pages — and how reliable is it? Four engines (ChatGPT, Google AI Overviews, Google AI Mode, Gemini), 10 identical reruns each, 8 hiring-software categories.
When AI recommends the “best” B2B software, a large share of what it cites is the vendors’ own marketing: across 1040 answers, 27% of cited sources were vendor-owned pages — and it ranges 4.4× by engine, from 53% on ChatGPT down to 12% on Google AI Overviews. The shortlist is also unstable: 40% of the vendors an engine names appear in half their identical reruns or fewer. The engines mostly name the same vendors (76% agreement) — but cite very different sources (20% shared) to do it.
share of cited sources that are the vendors' OWN pages — ChatGPT vs Google AI Overviews, a 4.4x gap. How independent an AI's software advice is depends heavily on which AI.
27% overall · citation-weighted
of the vendors AI names appear in half their identical reruns or fewer — a coin flip or worse. Ask again, get a different shortlist.
967 appearances · 10 reruns
of cited sources were shared between the engines (43% by the looser overlap measure). They name similar vendors but read different evidence.
ChatGPT, Google AI Mode, Gemini
- Measured
- 2026-07-21
- Engines
- ChatGPT · Google AI Overviews · Google AI Mode · Gemini
- Categories
- 8 · B2B hiring software
- Reruns / engine
- 10
- Answers
- 1040 of 1280 queries
- Raw data
- download CSV
Which AI is actually independent? The self-citation gap
Every engine leans on the vendors’ own pages to some degree — but not equally. Share of each engine’s cited sources that were vendor-owned (a higher bar means more of its “advice” is the seller’s own marketing):
ChatGPT builds 53% of its answers from vendors’ own pages — 4.4× Google AI Overviews (12%). Note Google AI Overviews triggered for only 36% of these buying questions (25% of all attempts), so it rests on fewer answers — shown, not hidden.
The shortlist is a coin flip — at any threshold
Instability isn’t an artifact of a strict cutoff. Ask an engine the same question 10 times; here is the share of vendor appearances that held at each bar — even “more than half the runs” leaves 35% below it:
One example of how far the engines diverge on a single product: in interview scheduling software, a widely-used vendor was named in 37/40 Google AI Mode runs but just 4/40 ChatGPT runs — same product, near-opposite verdicts, same day. (Named by category, not company — this is measurement, not a callout.)
Category by category
Each category, its shortlist instability and self-citation — each links its live tracker and raw runs:
| Category | Coin-flip vendors | Vendor-owned sources | Vendors named | Raw runs |
|---|---|---|---|---|
| Applicant tracking system (ATS) | 45% | 18% | 16 | JSON |
| Interview scheduling software | 38% | 36% | 14 | JSON |
| Employee onboarding software | 34% | 27% | 13 | JSON |
| Pre-employment assessment software | 45% | 26% | 15 | JSON |
| Recruiting CRM software | 43% | 18% | 12 | JSON |
| Recruiting software (talent acquisition suites) | 42% | 16% | 15 | JSON |
| Recruiting / staffing agency software | 38% | 35% | 13 | JSON |
| Video interviewing software | 36% | 38% | 12 | JSON |
“Coin-flip vendors” = share of this category’s vendor appearances that showed in ≤50% of reruns. Small per-category differences aren’t a ranking — with ~10 runs per cell the margins overlap; read the direction, not the decimals.
What it means — for buyers, and for vendors
For buyers: a single question to a single AI is a sample, not a verdict — 4 in 10 of the vendors it names are effectively coin flips, and much of the “advice” is assembled from the sellers’ own marketing pages. Which engine you ask changes what evidence you’re shown. Shortlisting from one AI answer sees a fraction of the picture.
For vendors: being named once proves little and absence in one check proves less — your real position is a distribution across engines and reruns. And because the engines lean so heavily on vendors’ own pages, the independent record (reviews, communities, third-party write-ups) is the movable ground. None of this is a quality ranking: we measure how often a product is named, never whether it is good.
Method, conflicts, and how to check us
We ran 32 real buyer prompts 10× on each of four engines, in fresh sessions on July 21, 2026, across 8 B2B hiring-software categories — 1280 queries, 1040 of which returned an answer. Product aliases and sub-brands are matched. Stability, agreement and self-citation are computed only where an engine actually answered; the three that reliably answer (ChatGPT, Google AI Mode, Gemini) carry the cross-engine figures, because Google AI Overviews triggered too rarely to demand agreement from it. Full method: how we check.
Our conflict, stated plainly: CrediGeo sells exactly the thing this study concludes you need — multi-engine, multi-run AI-visibility measurement. So don’t take our word for any of it. We published the exact prompts, every raw answer, and the aggregation script: download the summary CSV or each category’s raw run file and recompute the numbers yourself. Reproducible ≠ re-measurable: recomputing our published runs gives our figures exactly; re-asking the engines today gives different numbers — that drift is the finding.
Limits: a dated snapshot of a moving target. We report MENTIONS, not quality. Per-cell samples are ~10 runs, so treat single-point differences as directional (±several points), not precise. No vendor paid to appear or influence any figure. Corrections / right of reply: if an alias gap misrepresented a product, write to [email protected] — we re-check against the raw runs and correct publicly.
Cite this study
CrediGeo, State of AI Software Recommendations 2026. Measured July 21, 2026. https://credigeo.com/research/state-of-ai-software-recommendations/
Canonical stat: “53% of the sources ChatGPT cites for B2B-software recommendations are the vendors’ own pages, vs 12% for Google AI Overviews — and 40% of the vendors AI names appear in half their identical reruns or fewer (CrediGeo, 2026).”
Want to see where your product lands? Get a free run-by-run breakdown — your buyers’ real prompts on ChatGPT and Google’s AI Mode (the two most-used surfaces), run by run, with the sources each answer cited. It’s the same instrument behind this study. See a real sample report first →
Get your free breakdownQuestions about this study
- How much of AI’s software advice comes from the vendors themselves?
- In this July 21, 2026 study, 27% of all the sources AI engines cited to build B2B-software recommendations were the recommended vendors' own pages — and it varies 4.4x by engine: ChatGPT drew 53% of its cited sources from vendor-owned pages, Google AI Overviews just 12%.
- How reliable are AI software recommendations?
- Not very, and it's not a knife-edge: 40% of the vendors an engine names appear in half their identical reruns or fewer. Even at a strict "named in 8 of 10 runs" bar, only 48% of appearances are stable. A single AI answer is a sample, not a verdict.
- Do the AI engines recommend the same software?
- Mostly the same vendors — the ChatGPT, Google AI Mode, Gemini agreed on 76% of the vendors any of them named — but they read very different things to get there: they shared only 20% of their cited sources (43% by the looser overlap measure). Same shortlist, different evidence.
- Why should I trust a measurement company’s study on this?
- You shouldn't take our word for it — we profit if you believe this, so we published everything to check it: the exact prompts, every raw answer, and the script that computes the numbers. Download the data and recompute it yourself. That's the whole point.
Get the run-by-run breakdown for your category
We run your buyers’ questions across ChatGPT and Google’s AI Mode and send you who got named, run by run, with the exact prompts. Free, no call required.
That didn’t go through — please check the three required fields and send again.
Prefer to talk first? Book a live look (opens in a new tab) — your category on screen, no deck. Or read the real sample report.