Exactly how we measure your AI visibility
the runs below are real · measured 2026-07-16 · contract management software
We ask the questions your buyers ask, multiple engines, several runs each, and count who gets named. No blended score, no black box. This page is the whole method, and every number on it links to the raw answer it came from. Ask the engines yourself and the list will move. That movement is what we measure.
What exactly do we ask?
The questions a real buyer types when they’re shortlisting software in your category, things like “best [category] software,” “[competitor] alternatives,” “software for [job-to-be-done].” We build the set with you, in your buyers’ words. The wording then stays identical across every run and every month. Change the question and you can no longer compare the answers.
The prompt set is frozen. Because it could be gamed
A sophisticated buyer should ask this: couldn’t you pick prompts that flatter the story? Yes. Anyone selling a visibility number could. That’s why the prompt set is agreed with you up front, published inside your report, and never changes without a dated note in the methodology log. If a prompt is added or removed, you see when and why. A share-of-answer number without a fixed, visible prompt set isn’t a measurement; it’s an argument.
Why multiple engines. And never one blended number?
Because the engines read different sources and answer differently. We measure up to four surfaces, ChatGPT, Google AI Overviews (the AI answer on a normal Google search), Google AI Mode, and Gemini. And every measurement states exactly which engines it covered. In our July 16, 2026 category measurement, one major vendor was named in 12 of 20 ChatGPT runs and 0 of 12 Google AI Mode runs. A single “share of answer” number blends that into meaningless mush. We run each engine separately and report them side by side, the gap is usually the most useful thing on the page. One honest wrinkle: AI Overviews doesn’t appear for every query, so we also report how often it appeared at all. That trigger rate is itself a finding.
Why several runs each?
Because the answer moves. Here are five real, consecutive fresh-session ChatGPT runs of the same question, from the same morning. This is the actual log behind our sample report:
| Run | Vendors the engine named |
|---|---|
| 1 | Ironclad · LinkSquares · Juro · SpotDraft · PandaDoc · Gatekeeper · DocuSign CLM · Agiloft |
| 2 | Ironclad · Juro · LinkSquares · PandaDoc · Conga · DocuSign · Icertis |
| 3 | Ironclad · Juro · LinkSquares · Conga · DocuSign · Icertis · Agiloft |
| 4 | Agiloft · Conga · DocuSign · DocuSign CLM · Ironclad · Juro · PandaDoc · SpotDraft |
| 5 | Agiloft · DocuSign · DocuSign CLM · Ironclad · Juro · LinkSquares · PandaDoc · SpotDraft |
Count it yourself: Icertis appears in runs 2 and 3. And in none of the others. Gatekeeper appears once and vanishes. One run can make any vendor look present, or absent, by pure luck. This is also why we never say “ranked #1 in ChatGPT”: there is no stable ranking. We count how often you’re named across repeated checks, per engine, and date every figure.
What we report
Not a score, a distribution, and the context to act on it:
- Share of answer, per engine, how often you’re named across the runs on each engine, shown separately, never blended.
- Run-to-run stability, which prompts hold and which flip, so you know what’s a moat and what’s luck.
- The competitor leaderboard, who actually dominates the answers in your category.
- Cited sources, the domains the engines pull from to build those answers, the map for where authority is earned.
Audit it against the archive, then ask the engines yourself
Two different things, both yours. Audit. Every published number ships with its archived raw answers, full text, timestamps and cited URLs (the sample’s complete archive is public). An AI answer can’t be re-created later, so the archive is the audit trail. Prompts. You get the exact wording we used, so nothing about the question is hidden.
The exact question we asked
best contract management software for mid-market companies
That is the wording, verbatim. Every number we publish ships with the prompt that produced it and the archived raw answers behind it, so you can audit what we recorded line by line. Ask the engines yourself and the list will move. These systems vary by account, history, location and model version. That movement is the reason we run each question many times instead of once, and it is the thing we measure.
When is movement real?
Only once it clears run-to-run noise, a number that moved by one run hasn’t moved. Shown with real numbers in the sample report →
What we can’t measure. And won’t pretend to
Mentions, not revenue: anyone claiming to isolate “named more often” → “more pipeline” is selling attribution they don’t have. And no one can promise a spot in an answer. AI output moves daily and can’t be bought. What we control, we commit to: measurement monthly, work shipped and logged with dates.
Get your free breakdown
A run is one question, asked once, in a clean session. We ask your buyers’ questions on ChatGPT and on Google’s AI Mode, more than once on each, then send you who got named in every answer with the exact prompts.
That didn’t go through, please check the required fields and send again.
What arrives, exactly
- Engines
- ChatGPT and Google AI Mode
- Runs
- 3 ChatGPT and 2 AI Mode per question, each a fresh session
- You get
- Every answer, who was named in it, and the exact prompts
- Arrives
- Within 2 business days, from a person
- Cost
- None, and no call required
Prefer to talk first? Book a walkthrough (opens in a new tab). Your category on screen, no deck. Or read the real sample report.