AEO summary: AI answer engines do not recommend products by brand size. We asked ChatGPT and Gemini the same buyer question 96 times across 6 software categories. Each category had 2–3 near-permanent "default" brands (named ~100% of the time), the #1 position still flipped in 83% of categories, the middle tier swung wildly run to run, and several household names (Salesforce for SMB CRM, LastPass, ConvertKit, Basecamp) barely appeared. The takeaway: a single check tells you almost nothing — appearance rate over repeated, dated observations is the honest metric.
Why this matters
Founders and marketers increasingly ask, "Does ChatGPT recommend us?" — then run the query once, see (or miss) their brand, and draw a conclusion. That conclusion is usually wrong, because a single AI answer is one sample from a non-deterministic system.
We wanted real data on how (in)consistent those recommendations actually are — and whether being a well-known brand is enough to get named.
Methodology
- Categories: 6 — project management software, CRM software, email marketing software, accounting software, help desk software, password managers.
- Prompts per category: 1 buyer-intent prompt (e.g. "What is the best CRM for a small B2B sales team?").
- Engines: ChatGPT and Gemini, via provider-sampled API runs, with response caching disabled so every run is a fresh sample.
- Runs: 8 per engine → 16 runs per category, 96 observations total. All 96 returned successfully.
- Window: August 2026.
- What we recorded: for a curated set of well-known brands per category, whether each was named and in what order.
- Two metrics: appearance rate (share of runs naming a brand) and top-3 churn (how much the top-3 set changes between consecutive runs).
- Limitations: this is a pilot — one prompt per category, brand detection limited to a curated list, and results reflect the study window. AI answers shift over time and by region.
Headline findings
| Metric | Result |
|---|---|
| Median top-3 churn between runs | 0% |
| Categories where the #1 brand changed at least once | 83% |
| Median appearance rate of the most-named brand | 100% |
| Brands appearing in only one run (share of tracked) | 3% |
Read together, these say something specific: the top of each list is stable, but the ranking within it and everything below it is not.
Finding 1 — Every category has 2–3 near-permanent defaults
In all six categories, the leading brands appeared in 100% of runs. AI answer engines clearly have durable "default" recommendations they reach for almost every time.
Finding 2 — But the #1 slot is not fixed
Even with a stable top set, the brand named first flipped at least once in 5 of 6 categories (e.g. Monday.com ⇄ ClickUp). Position — which most people read as "the best" — is not deterministic.
Finding 3 — The middle tier is volatile
Below the defaults, appearance rates swung widely: Salesforce 38%, Mailchimp 81%, Brevo 63%, ActiveCampaign 31%, Trello 63%, Jira 38%, Dashlane 50%, NordPass 44%. This is where most brands live — and where a single check is most misleading.
Finding 4 — Brand size ≠ recommendation
The most striking result: several household names were effectively invisible for these buyer questions.
- Salesforce appeared in only 38% of "best CRM for a small B2B sales team" answers.
- LastPass, ConvertKit, Constant Contact, Basecamp, Wrike, Sage, and Copper each appeared 0% of the time in their categories.
Recognition does not automatically transfer to AI recommendations.
By category
| Category | Most-named brand | Its appearance rate | Top-3 churn | #1 changed? |
|---|---|---|---|---|
| Accounting software | QuickBooks | 100% | 0% | yes |
| CRM software | HubSpot | 100% | 0% | yes |
| Email marketing software | Klaviyo | 100% | 0% | yes |
| Help desk software | Zendesk | 100% | 0% | no |
| Password managers | 1Password | 100% | 0% | yes |
| Project management software | Monday.com | 100% | 0% | yes |
Full appearance rates
Project management software — Monday.com 100%, ClickUp 100%, Asana 94%, Notion 81%, Trello 63%, Jira 38%, Basecamp 0%, Wrike 0%
CRM software — HubSpot 100%, Pipedrive 100%, Zoho CRM 100%, Close 88%, Salesforce 38%, Freshsales 6%, Copper 0%
Email marketing software — Klaviyo 100%, Omnisend 100%, Mailchimp 81%, Brevo 63%, ActiveCampaign 31%, ConvertKit 0%, Constant Contact 0%
Accounting software — QuickBooks 100%, Xero 100%, FreshBooks 100%, Wave 100%, Zoho Books 94%, Sage 0%
Help desk software — Zendesk 100%, Intercom 100%, Freshdesk 63%, Help Scout 63%, Zoho Desk 38%, Gorgias 0%
Password managers — 1Password 100%, Bitwarden 100%, Keeper 100%, Dashlane 50%, NordPass 44%, LastPass 0%
What this means for your brand
- If you're a default (100%): your job is to defend it — the position isn't guaranteed, and the #1 slot moves.
- If you're in the volatile middle (30–80%): one check is not evidence. Your real standing is your appearance rate across repeated, dated observations.
- If you're a big brand that's invisible (0%): brand recognition didn't carry over. AI leans on the sources it trusts — comparison pages, review platforms, and well-structured content — not on how famous you are.
How we'd track your brand
mentionpop runs this exact kind of repeated sampling for your category and market, then reports your appearance rate, cited sources, competitor appearances, and the first corrective move. Run a free check or read the sampled observation methodology.
This is a pilot dataset (6 categories, 96 observations). We'll expand prompt coverage and re-run it periodically to track how AI defaults shift over time.
Frequently asked questions
How was consistency measured?+
For each category we asked one buyer-intent prompt 8 times on ChatGPT and 8 times on Gemini (16 runs per category, 96 total), with response caching disabled so each run was a fresh sample. For a curated set of well-known brands per category, we recorded whether each was named and in what order, then computed appearance rate (share of runs naming it) and how much the top set changed between runs.
Does a high appearance rate guarantee recommendations?+
No. AI answers are non-deterministic. A high appearance rate means a brand shows up consistently in sampling, not that every user in every session will see it. It is evidence of visibility, not a guaranteed ranking.
Isn't the whole thing just random?+
No — and that's the interesting part. The top 2–3 brands per category were extremely stable (often named in 100% of runs). The volatility is concentrated in the middle tier and in which brand takes the #1 slot. So it's neither fixed nor random: there are durable defaults plus a shifting middle.