+ Blog

How to Get Your Startup Cited by ChatGPT, Claude and Google AI Overviews

Kara Silverman
15 min
September 23, 2026
How to Get Your Startup Cited by ChatGPT, Claude and Google AI Overviews

TL;DR

  • 3.0 percent of relevant Google AI Overviews include the median enterprise B2B brand, across 45 million keywords and 828 companies. When a buyer asks an AI assistant for a shortlist and you are not in the answer, you were never in the running.
  • Being cited and ranking are different jobs. A ranking puts your page on a results page. A citation means an engine chose your page, or somebody else's page about you, to support the answer it wrote.
  • Most citations go to sources you do not own. Muck Rack's May 2026 sample of 25 million links from ChatGPT, Claude and Gemini answers put 84 percent of citations on earned media, and Reddit dominated citations in Semrush's July 2025 study. Your own site is one of four source types, and never the whole answer.
  • Nobody can guarantee you a citation, and we quote the engines saying so below. What you can do is run the six-step method in this guide, then measure it every month with free tools from Bing, Google and OpenAI.

Why do your competitors show up in ChatGPT and you don't?

Because the engine found more evidence for them. That is the whole answer, and the rest of this guide is about what counts as evidence and where it comes from.

Two words need separating first. A mention is the engine naming your company in its answer. A citation is the engine pointing at a specific page as a source: a link, a title, a quoted passage. You can get one without the other. ChatGPT can recommend you while citing an analyst's comparison rather than your site. An AI Overview can cite your documentation without saying your name in the summary. Mentions are what the buyer sees. Citations are how the engine got there, and they are the part you can work on.

You will also meet two acronyms. GEO, generative engine optimisation, is the work of getting your brand and pages into AI-written answers. AEO, answer engine optimisation, is structuring information so an answer engine can use it. We call the whole thing AI visibility, and we treat it as the score of everything you already do rather than a new channel. When someone asks ChatGPT to recommend tools in your category, the answer reflects your entire footprint: earned media, owned content, site structure, community presence. That is our position, not a citation. The rest of this piece is the evidence behind it.

How do you get cited by ChatGPT?

Start with the crawler. OpenAI's publishers and developers FAQ says that for your content to be included in ChatGPT's summaries and snippets you must not block OAI-SearchBot, and that this may mean editing robots.txt. Check your CDN and bot protection too. A firewall that challenges unknown crawlers shuts out OAI-SearchBot as surely as a disallow line does. Blocking OAI-SearchBot keeps a page out of ChatGPT's summaries and snippets. OpenAI adds that if it gets the URL from somewhere else it may still surface the link and page title in its Atlas browser, and that a noindex tag stops even that, provided the crawler is allowed in to read the tag.

Search access and training are separate switches. GPTBot is the crawler that governs training. OAI-SearchBot governs search. You can allow one and disallow the other, and plenty of companies do.

That is the whole of what OpenAI documents for publishers: let the search crawler in, or block it, plus noindex, to stay out. There is no ranking formula in the FAQ and no submission form. What the FAQ does hand you is a way to see results. ChatGPT adds utm_source=chatgpt.com to every referral link, so visits that began in a ChatGPT answer show up in Google Analytics under their own source. Hold that thought for the measurement section.

How do you show up in Claude?

Anthropic documents three crawlers, and founders mix them up. According to Anthropic's crawler page, ClaudeBot "helps enhance the utility and safety of our generative AI models by collecting web content that could potentially contribute to their training". Blocking it is a training decision. It says nothing about whether Claude can read your site when a user asks a question today.

That job belongs to the other two. Claude-User fetches pages when "individuals ask questions to Claude", and Anthropic says disabling it "may reduce your site's visibility". Claude-SearchBot "navigates the web to improve search result quality for users", and disabling it "may reduce your site's visibility and accuracy in user search results". All three honour robots.txt. So the first check is the same as for ChatGPT: open your robots.txt and make sure a rule written for ClaudeBot two years ago is not also shutting out Claude-User and Claude-SearchBot.

When Claude searches, it shows its sources. Anthropic's web search documentation says "citations are always enabled for web search", and each citation carries the source URL, the page title and the passage cited. That matters for the buyer's-eye view. A founder who asks Claude for a shortlist sees exactly which pages produced each name.

Two more things. Anthropic publishes no ranking formula either. And Claude is not a footnote in the data: Muck Rack's May 2026 sample was drawn from ChatGPT, Claude and Gemini answers, so the earned-media findings below apply to it directly.

How do Google AI Overviews, AI Mode and Copilot choose what to cite?

Google: it is still Search

Google is the only engine that has written a guide. Optimizing your website for generative AI features on Google Search, updated in July 2026, says AI Overviews and AI Mode are "rooted in our core Search ranking and quality systems". Two mechanisms are named. Retrieval-augmented generation, which Google also calls grounding, uses the core ranking systems to pull current pages from the Search index and then reads them to write the answer. Query fan-out sends "a set of concurrent, related queries" to fetch more results. Google's own example turns "how to fix a lawn that's full of weeds" into searches for herbicides, chemical-free removal and prevention.

The eligibility rules are short. A page must be indexed and eligible to show a snippet, and Google's AI features page says there are "no additional requirements". Inclusion in AI Overviews and AI Mode is on by default for every property, but it can be switched off in Search Console's Search generative AI control, so check that nobody on your team has turned yours off. Then comes the line every vendor should be made to quote: "Indexing and serving aren't guaranteed."

The guide also lists what you can ignore, under its own "Mythbusting" heading. You do not need an LLMs.txt file or "other 'special' markup", because "Google Search itself doesn't use them". You do not need to chunk content into fragments. You do not need to rewrite anything just for AI systems. Structured data "isn't required for generative AI search", though it still earns rich results. And "seeking inauthentic 'mentions' across the web isn't as helpful as it might seem". Google's summary of all this is blunt: "optimizing for generative AI search is optimizing for the search experience, and thus still SEO", and if somebody is selling you AEO or GEO services, it points you to its guidance on evaluating third-party SEO advice.

What Google does want is "valuable, non-commodity content" with "a unique point of view". Its example: "7 Tips for First-Time Homebuyers" is commodity, while "Why We Waived the Inspection & Saved Money: A Look Inside the Sewer Line" is not, because only somebody who did it could have written it.

Copilot: the most specific controls of any engine

Microsoft's Bing Webmaster Guidelines describe how Bing "discovers, crawls, indexes, evaluates, and surfaces content across Bing search experiences, Copilot, and grounding API results", one foundation for all three. The controls are precise. NOARCHIVE "prevents content from being used in Copilot responses and grounding results". NOCACHE "limits Copilot to using only the URL, title, and snippet, reducing citation depth and answer quality". IndexNow tells Bing the moment a URL changes. Two of the guidelines read as if they were written for this article: "Define entities clearly and consistently", because "clear entity definition improves grounding visibility and citation accuracy", and "focus each URL on a single topic", because such pages "are more likely to be selected for grounding results".

What do the citation studies actually show?

Four studies, each with a date and a sample, and each saying something narrower than its headline.

Earned media wins. Muck Rack's What is AI reading?, May 2026 edition, analysed more than 25 million links from ChatGPT, Claude and Gemini answers across 17 industries. Earned media accounted for 84 percent of citations, journalism alone for 27 percent, and paid or advertorial content for 0.3 percent. Across the three editions since July 2025, the earned share has stayed between 82 and 89 percent. That is a share of what the engines cite, not of what trained them.

Reddit dominates. Semrush's AI Mode comparison study of July 2025 looked at more than 150,000 citations and found Reddit "a leading citation source across all LLMs we studied". Foundation's study of 50 B2B brands put Reddit at 20.8 percent of external citations overall and 30.9 percent on unbranded discovery queries, the ones where the buyer has not yet named a vendor.

Human writing gets cited. Graphite's October 2025 study of 31,493 keywords across 10 categories found that 82 percent of articles cited by ChatGPT and Perplexity were human-written, against 86 percent of Google-ranking articles. It does not prove that authorship causes selection. It does say a content farm is a poor bet.

Most B2B brands are absent. Walker Sands' benchmark of 45 million keywords across 828 enterprise B2B companies found the median brand in 3.0 percent of relevant AI Overviews. That is Google only. We have not found a published equivalent for ChatGPT or Claude.

The studies agree on where citations come from, and it is mostly not the brand's own site. We group those sources into four buckets. The method below works through all four, though not one step per bucket: the owned and reference work sits inside step 2, editorial is step 3, and community is step 5.

Source bucket What it holds Why it matters
Owned Your site, docs, blog Necessary for the facts, rarely sufficient alone
Trusted editorial Journalism, trade press, analysts Where somebody other than you says what you say about yourself
Community and social Reddit, forums, professional groups Where an engine finds experience a product page cannot manufacture
Reference and aggregator Directories, review sites, comparison pages Where engines confirm your category and your relationships

What do you do first? The method, in order

1. Map the prompts buyers ask

Write down the questions a buyer would type before they know your name: shortlist prompts, comparison prompts, "alternatives to", "how do teams handle", and the risk questions your sales calls always reach. Run each one in ChatGPT, Claude, Google AI Mode and Copilot. For every answer, record who was named, which pages were cited, and the date. This is your baseline and your target list at once, because the cited pages tell you which outlets, communities and reference sites the engines already trust in your category. We baseline 300 to 500 prompts for clients. A founder with a spreadsheet can start with 30.

2. Fix the entity

Every engine section above came back to the same thing: the engine has to know what you are. Bing says it outright, and Google's grounding runs on the same index as Search. So give them one name, one category sentence, and the same facts everywhere: your site, LinkedIn, Crunchbase, review profiles, and the bio your founder uses when quoted. Add Organization schema that matches the visible page. Google says schema is not required for AI features, so treat it as clarity rather than a lever. Part 2 of the Althea Score, on back-end infrastructure, walks through the technical checks.

3. Earn coverage where the engines already cite

Take the cited pages from step 1 and count the domains. The publications that appear again and again are your media list, and they are worth more to you than a bigger outlet the engines never cite in your category. Pitch them something they can report: original data, a documented change in the market, a practitioner's account of a decision that went wrong. A funding announcement gets a paragraph. A finding gets a story, and stories get cited. If you want the press-release version of this argument, we wrote it.

4. Make your own pages worth citing

Google's guide and an independent measurement study agree on what a citable page looks like, from opposite directions. Google asks for a unique point of view and non-commodity content. The authors of an arXiv study from April 2026 found that the pages with the most influence on the answers "tend to be longer, more structured, semantically aligned, and richer in extractable evidence such as definitions, numerical facts, comparisons, and procedural steps". They measured what those pages have in common rather than what made the engines pick them, which is still the most useful description of a citable page we have seen. So make specific claims, show the data or the example behind each one, name the author, date the facts, keep one topic per URL, and structure the page for a reader who is skimming. Do not chop it into fragments or write a robotic answer under every heading. Google says you do not need to, and buyers hate it.

5. Seed human proof

Reddit's share of citations is not an invitation to fake it. Google names "inauthentic mentions" as a thing to skip, and unbranded queries, where Reddit's share is highest, are exactly where a planted thread gets caught. Show up where your buyers already ask each other for advice, answer the specific question with real detail, say who you are, and mention the product only when it solves the problem in the thread. Encourage customers to describe their setups in their own words. Part 4 of the Althea Score covers how the engines read community consensus.

6. Track the same prompts monthly

Re-run step 1 with the same prompts every month. Log mentions and citations separately, note which competitors appear, and check whether the description of you is right. The pattern tells you the next move. Missing citations usually mean thin pages or no third-party coverage. Wrong descriptions usually mean inconsistent entity facts. Change one thing at a time so you can see what moved.

How do you measure whether it's working, without paying for tools?

Four free instruments, each measuring a different thing. Keep them separate.

  1. Bing Webmaster Tools, AI Performance report. In public preview since February 2026, the report shows "which pages are cited, how visibility trends change over time, and the grounding queries associated with your content" across Copilot and partner experiences. Read the counts as Bing tells you to. They reflect "how often pages are cited, not their importance, ranking, or role within a response", the data "represents a sample", and a citation "does not represent traffic, clicks, or user engagement". We have not found another free report that shows you grounding queries.

  2. Google Search Console, Generative AI performance report. The report shows impressions from AI Overviews and AI Mode over time, by page, device and country. Google says it finished rolling the report out to all sites worldwide on 31 August 2026. If you cannot see it, the usual reason is too few impressions to report, or somebody excluded the site through the Search generative AI control. It measures impressions rather than citations, and it will not tell you why a page was chosen.

  3. Google Analytics 4, ChatGPT referrals. Because OpenAI stamps utm_source=chatgpt.com on referral links, a source filter in your acquisition reports isolates visits that began in a ChatGPT answer. Visits only. An answer that named you without a click leaves no trace here.

  4. Your tracked prompt set. Step 6 of the method. It is the only one of the four that covers Claude, and the only one that shows you the competitors.

No free tool we have found reports Claude citations natively, and every engine's answers vary between runs, so treat a single month as noise and three months as a trend.

What can nobody promise you?

A citation. Google's page on hiring an SEO says "No one can guarantee a #1 ranking on Google", its AI guide says "Indexing and serving aren't guaranteed", and the same guide adds that "No third-party tool has access to our internal ranking or AI systems". OpenAI and Anthropic document their crawlers and nothing more. A proposal that guarantees a mention in ChatGPT is guaranteeing something its author cannot see.

There is a subtler reason to distrust guarantees. The arXiv study above analysed 602 controlled prompts and 21,143 citations across ChatGPT, Google and Perplexity, and it separated citation selection (the engine linked to you) from citation absorption (your page actually shaped the answer). The two diverge. Perplexity and Google cite more sources per answer. ChatGPT cites fewer but leans on them harder. Breadth and depth are not the same thing, so effects vary by engine and by question, and anyone quoting a single win as proof of a method is quoting an anecdote.

What a credible partner can promise is the work: crawler access checked, entity fixed, coverage earned in the outlets that get cited, pages built to be absorbed, community proof that is real, and a monthly measure of whether any of it moved. That is what we sell, and it is all anyone can honestly sell.

FAQs

How long before you get cited by ChatGPT or Claude? Neither company publishes a timeline and neither guarantees inclusion. In our work the baseline lands in the first month of tracking and early movement shows within 60 to 90 days, with the gains compounding after that. Crawler access and entity facts move fastest. Earned coverage takes longest and lasts longest.

Does schema markup help with AI citations? It helps engines identify you, and Bing says clear entity definition "improves grounding visibility and citation accuracy". Google says structured data "isn't required for generative AI search". Use it for clarity and rich results, and do not expect it to buy a citation.

Can an agency guarantee you a mention? No, and the section above gives the platforms' own words. Ask instead what they will change, in which source bucket, and how they will measure it monthly.

Do you need to block or allow AI crawlers? Allow the search crawlers if you want to be cited: OAI-SearchBot for ChatGPT, Claude-User and Claude-SearchBot for Claude, and Bing's crawler for Copilot, which grounds on Bing's index. The training crawlers, GPTBot and ClaudeBot, are a separate decision you can make either way.

Does Perplexity matter for a B2B startup? Put it in your prompt set if your buyers use it. It cites generously and shows its sources, which makes it a useful window on what the other engines are reading. Fewer of the founders we talk to use it as a first stop, which is why it sits behind ChatGPT, Claude and Google here.

Should you write differently for AI? Google says no: "You don't need to write in a specific way just for generative AI search." Write for the buyer who is skimming, and make the facts easy to find. The arXiv finding on absorbed pages is the same advice from the data side.

How does PR Plus do this, and where do you start?

PR Plus is our name for running that method as one programme instead of four vendors. Three of its four components sit on a source bucket each. Authority is traditional PR, aimed at the editorial sources the engines cite. Infrastructure is SEO and schema, which covers the owned pages and the entity facts the directories and review sites repeat. Human Proof is community. The fourth, The Bridge, is the content that carries one set of claims across all of them.

The engagement itself runs in five steps of its own, in this order. Audit: we review your SEO content and technical setup, your PR footprint across owned and earned media, and your AI visibility and entity baseline, and we hand you a Brand Presence Report with the fixes ranked. ICP and Messaging: we agree who you are selling to and the one sentence about you that every source should carry. Visibility Tracking: we baseline 300 to 500 prompts drawn from real conversations, then re-run a tracked set monthly. Strategy: we choose which outlets, pages, fixes and communities come first. Execution: we do the PR, the content, the infrastructure and the community work, and we report against the baseline every month.

You do not have to wait for Google, OpenAI or Anthropic to publish a ranking formula to find out where you stand. Start with a Brand Intelligence Audit: five days, a deck and an hour with us, and you leave with your Althea Score benchmarks and the fixes for the next 30 to 45 days. If you want to see how that connects to ongoing work, here is what we do.