
TL;DR
- 84% of the links cited by ChatGPT, Claude, and Gemini are earned media, not owned pages (Muck Rack, May 2026). Page structure is one input among many.
- Our hub defines GEO and AEO: GEO (Generative Engine Optimization) targets LLM-powered platforms like ChatGPT and Perplexity; AEO (Answer Engine Optimization) is structuring a page so an answer engine can lift a direct answer from it. This post is about the second job.
- Google Search does not use llms.txt, does not require chunking or special markup, and stopped showing FAQ rich results on 7 May 2026. Advice written before those changes is still in circulation, FAQ schema recommendations included.
- Google asks for clear headings, crawlable content, and a page indexed and snippet-eligible. Bing adds two rules of its own: one primary topic per URL, and essential information near the top. In plain words: Google wants a page a reader can follow and an engine can index. Bing wants one question per page, answered at the top.
What do AI answers cite before page structure matters?
When Muck Rack’s Generative Pulse team analyzed more than 25 million links from ChatGPT, Claude and Gemini responses across 17 industries in May 2026, earned media accounted for 84% of citations. Journalism alone made up 27%. Paid and advertorial content made up 0.3%. Our read of that split is simple: a blog restructure cannot create the independent source signals earned media provides, so the structure question comes second to the coverage question.
Brands are mostly absent from these answers whatever their pages look like. Walker Sands found that 3.0% of relevant Google AI Overviews include the median enterprise B2B brand, across more than 45 million keywords, 828 companies and 14 industries. That benchmark covers enterprise brands in Google AI Overviews only, so it says nothing about startups or about ChatGPT, and we would not stretch it further.
Community discussion is the other input. Foundation found that Reddit accounted for 20.8% of citations to the top 50 external domains across 50 B2B brands, rising to 30.9% on unbranded discovery queries. Reddit threads hold independent discussion that a company’s own pages cannot supply. That is our read of why they get cited, not something Foundation measured.
So where does structure fit? It decides how easily an engine can extract a definition, a fact or a comparison once it has found your page. Our approach puts extractable owned content alongside earned media and search signals. The rest of this post is about the smaller part of that work, the part nobody else has to agree to.
What does Google say about page structure for AI Overviews and AI Mode?
Everything Google says here applies to Google Search, meaning AI Overviews and AI Mode. Google does not document how ChatGPT, Claude, Perplexity or Copilot choose sources, and nothing below should be read as if it did.
Google’s guide to generative AI features says those features are “rooted in our core Search ranking and quality systems,” which is why its structural advice is the familiar kind: “Write content for your human audience and make sure the content is well written and easy to follow. People generally appreciate it when web pages are organized by paragraphs and sections, along with headings that provide a clear structure.” Google ties structure to readability, not to a guaranteed place in an answer.
Eligibility is the ordinary search bar. Google’s AI features page states: “To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements. There are no additional technical requirements.” Google’s newer generative AI guide adds one condition that page leaves out: “a site must be included in Search generative AI features in Search Console to be eligible for display in generative AI features on Google Search.” That setting is on by default, so it is something to confirm nobody has turned off, not something to turn on.
The AI features page also lists the fundamentals it still calls worthwhile: crawling allowed in robots.txt and by your CDN, content findable through internal links, a good page experience, important content available as text, and structured data that matches the visible text on the page.
Perfectly semantic markup is not required, because “the web in general is not valid HTML, and Google can understand it,” though semantic HTML “helps other types of users, such as screen readers.” On length, the guide is blunt: “There’s no ideal page length, and in the end, make pages for your audience, not just for generative AI search.”
Google also answers the question every AEO vendor gets asked. Its featured snippet documentation says: “How can I mark my page as a featured snippet? You can’t.” You can make a passage easy to lift. You cannot nominate it.
Which AEO tactics does Google say you can ignore?
Google lists five common AEO recommendations it says you can ignore for Google Search, including AI Overviews and AI Mode. In its own words:
- llms.txt and special markup. “You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn’t use them.”
- Chunking content. “There’s no requirement to break your content into tiny pieces for AI to better understand it. Google systems are able to understand the nuance of multiple topics on a page and show the relevant piece to users.”
- Rewriting content for AI. “You don’t need to write in a specific way just for generative AI search.” Google’s systems understand synonyms and general meaning, so you do not need a page for every phrasing of a question. Building one anyway, “primarily to manipulate rankings or generative AI responses,” violates the scaled content abuse spam policy.
- Manufactured mentions. “Seeking inauthentic ‘mentions’ across the web isn’t as helpful as it might seem.” Google says its core ranking systems focus on high-quality content while other systems block spam, and its generative features depend on both. That is Google backing the point this post opened with: owned-surface volume does not substitute for coverage someone chose to publish.
- Overfocusing on structured data. “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add.” Google still recommends it as part of ordinary SEO, because it keeps a page eligible for rich results.
The guide’s summary line is “Prioritize effective SEO strategies over ‘AEO/GEO hacks’.” For Google Search, that means indexing, crawl access, internal links and page quality before any AI-specific formatting.
On llms.txt specifically, Google added a clarification on 15 June 2026: the file “won’t negatively or positively impact your visibility or rankings,” and you are free to maintain one for other services that use it. We publish an llms.txt on althealabs.co ourselves for the services that do read it, and we expect nothing from Google for it.
Should you still add FAQ schema?
Google stopped showing FAQ rich results on 7 May 2026 and removed the feature’s documentation in June, so FAQ markup no longer produces that search result. That is the fact. Our position sits next to it: we still keep FAQ markup where the questions are real, for the same reason we keep Organization and Article markup, as hygiene that removes ambiguity for the engines that read it. What we no longer do is add questions to a page in order to carry the markup.
The visible question-and-answer section is a different matter. In our experience a direct question followed by a direct answer is the easiest thing for an AI answer to lift, and Bing’s AI Performance guidance names FAQ-style content among the formats that make information easier to reference in its answers. Keep the section.
Bing’s structured data page says annotations tell its crawlers what kind of content a page holds, but the presence of markup alone does not guarantee a rich snippet. That page says nothing about Copilot or grounding.
What does Bing say about structure for Copilot answers?
Bing publishes the most specific page-structure guidance any engine offers, and it covers Bing search experiences, Copilot, and grounding API results. Six of its guidelines bear on structure:
- Structure content clearly using HTML. A logical H1 through H6 hierarchy, semantic elements, and accurate title tags and meta descriptions. Missing, duplicate or overly short titles and descriptions “may reduce indexing reliability, ranking, and eligibility for grounding results and citations.”
- Use structured data accurately. Bing says structured data “may support clearer grounding but does not guarantee visibility or grounding traffic,” and that markup “must accurately reflect visible content.”
- Ensure content can be verified independently. URLs are more likely to be selected for grounding queries and citations when facts and definitions are explicit, key statements do not rely on implied content, and important information is visible on the URL itself.
- Define entities clearly and consistently. Clear, consistent naming for people, organizations, products and locations “improves grounding visibility and citation accuracy.”
- Focus each URL on a single topic. “URLs focused on a primary topic are more likely to be selected for grounding results.”
- Surface key information early. Place essential information near the top; Bing says early clarity improves grounding visibility and reliability.
Bing’s AI Performance page adds format advice: “descriptive headings, concise sections, tables, and FAQ-style content,” plus “examples, data, and cited sources,” matched to query intent, such as comparison tables for comparison queries. It notes that these insights “do not guarantee citation outcomes.”
Two robots directives matter here. Bing says NOARCHIVE “prevents content from being used in Copilot responses and grounding results,” and NOCACHE “limits Copilot to using only the URL, title, and snippet.” A template carrying either tag by accident will not be rescued by restructuring.
Bing frames all of this as eligibility, with a caveat in its own words: “SEO does not guarantee rankings or traffic, and GEO does not guarantee grounding or citations in AI experiences.”
What do OpenAI, Anthropic and Perplexity document?
OpenAI publishes crawler controls, not formatting guidance. Its bots page says “OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features,” and that sites opted out of it “will not be shown in ChatGPT search answers, though can still appear as navigational links.” GPTBot crawls content that may be used for training. ChatGPT-User fetches pages when a user asks and is not used to decide whether content may appear in search.
Anthropic documents three crawlers: ClaudeBot for content that could contribute to training, Claude-User for pages fetched at a user’s request, and Claude-SearchBot, which analyzes content “to enhance the relevance and accuracy of search responses.” Anthropic says disabling the last two “may reduce your site’s visibility.”
Perplexity follows the same pattern. PerplexityBot “is designed to surface and link websites in search results on Perplexity” and is not used to collect content for training. Perplexity-User supports actions a user takes inside the product.
As far as we can find, none of the three tells site owners how to structure headings, where to place an answer, how long a page should be, when to use a table, what schema to add, or whether to publish an llms.txt. Any claim about the page format ChatGPT, Claude or Perplexity prefers rests on testing or inference, not on anything those companies publish. When a proposal cites “what ChatGPT wants,” the honest version is “what we observed.”
How should you structure a page for answer engine optimization (AEO)?
The short answer: one question per page, the answer in the first paragraph, headings that say what each section answers, facts and definitions stated outright, a table only for things that compare, honest titles and descriptions, and only the markup that describes what a reader can see. The seven sections below say whose rule each one is. We covered the same practices more briefly in GEO, AEO, and SEO; here we name only what Google and Bing put in writing, and we say so when neither does. This post is built to the same rules, so you can check them against what you are reading.
One page, one question
Give each URL one primary question, then the subquestions needed to answer it. This is Bing’s rule. Google does not ask for it, since its systems can read several topics on one page, but it does warn that building separate pages for every query variation, primarily to manipulate rankings or generative AI responses, is scaled content abuse.
Put the answer first
Answer the primary question near the top, before background or company positioning. Bing’s rule again. Google says nothing about placement for its AI features, so with Google this is a courtesy to readers rather than a documented lever. We do it anyway: a reader who gets the answer in the first paragraph stays for the evidence.
Write headings a reader would ask
Use headings that name the question each section answers. Both Google and Bing ask for clear headings, a structural point they share. Headings should help a reader follow the argument rather than repeat keyword variants.
Make facts and definitions explicit
State the definition, the number and the source on the page, in sentences that keep their meaning when lifted out. This is Bing’s rule, not Google’s. A descriptive study of 21,143 search-layer citations found that high-influence pages tended to be “longer, more structured, semantically aligned, and richer in extractable evidence such as definitions, numerical facts, comparisons, and procedural steps.” That is an association, not proof that adding length or structure causes a citation.
Use a table only when the content is a table
Use a table when readers need to compare equivalent attributes across items, and a list when sequence or membership carries meaning. Bing’s AI Performance guidance names tables among the formats it suggests; Google’s guide does not mention them. Converting ordinary prose into blocks earns nothing from either engine.
Keep titles, descriptions, dates and bylines accurate
Keep the title and meta description specific to the page, and show real dates and bylines where readers expect them. Bing documents accurate titles and descriptions as inputs to indexing and grounding eligibility. Google strongly encourages “adding accurate authorship information, such as bylines to content where readers might expect it.” Visible publication and update dates help a reader judge currency; that is our judgment, not Google’s. A front-end SEO audit covers the titles and descriptions that contradict each other across a site.
Add schema that matches the page, and stop there
Add structured data only where a supported schema type describes content a reader can see. Both Google and Bing ask that markup match what is visible. Validate it, fix crawl and indexing problems first, and do not add fields that extend beyond the page. A back-end infrastructure audit separates schema problems from crawler and rendering problems.
Where do Google and the AI engines agree, and where do they differ?
Read the overlap as the baseline and the differences as the test of any vendor’s claims. If a proposal credits Google with Bing’s rules, or credits ChatGPT with rules nobody published, it was written from a template.
How do you tell whether a structure change did anything?
We have yet to see an AEO report that can do more than show visibility changed. None of them can show a heading rewrite caused it.
Google’s Search Console generative AI performance report counts impressions, meaning the times a link to your site was shown in AI Overviews or AI Mode. It covers those two Google features only. Compare impressions for the changed page group before and after publication.
Bing’s AI Performance report shows which pages were cited, and for which grounding queries, across Microsoft Copilot and partner experiences. Bing describes the data as a sample and says a citation “does not represent traffic, clicks, or user engagement.”
Track a fixed prompt set in Gauge, the AI visibility measurement platform we use with clients. Gauge samples ChatGPT, Perplexity, Gemini, Copilot, Google AI Overviews and AI Mode. It does not sample Claude. Keep the prompts and the scoring method the same across periods.
Change one template or page group at a time, record the date, and keep a similar group unchanged as the comparison. Do not launch a sitewide rewrite in the same month as a PR push. A controlled rollout is the closest thing to evidence this work can produce.
What should you ask an agency selling an AEO rewrite, and can you do it in-house?
An agency should diagnose your visibility problem before it prescribes a sitewide rewrite. Five questions tell you whether its recommendations come from engine documentation:
- Which engine documents each proposed change, and what does the source actually say?
- Google stopped showing FAQ rich results on 7 May 2026. What is your FAQ recommendation now, and why?
- Which pages will you change first, and which unchanged pages will be the comparison?
- How will you report Google impressions, Bing citations, referral traffic and tracked prompts separately?
- What result would tell you the rewrite failed?
A single marketing lead can do most of this in-house: tighten headings, move the answer up, name entities consistently, correct titles and descriptions, and add markup that matches the page. What you also need is analytics access to record a baseline, and the discipline to change one group at a time.
What structure cannot do is create third-party authority or community discussion. When those source signals are missing, an editorial rewrite works on the wrong layer. Althea Labs’ PR Plus works on those signals through its four components: Authority (traditional PR), Infrastructure (SEO and schema), Human Proof (community), and The Bridge (content). The pages this post describes are The Bridge; they do not replace the other three.
Before hiring, read what GEO agencies do and ask any agency to map every deliverable to a documented problem. Our view: an agency that cannot name the source behind a recommendation is selling a template under a newer label.
FAQs
How do GEO and AEO differ? As our hub puts it, GEO (Generative Engine Optimization) targets LLM-powered platforms like ChatGPT and Perplexity, while AEO (Answer Engine Optimization) is structuring a page so an answer engine, from Google’s featured snippets and voice assistants to ChatGPT, can lift a direct answer from it. The terms converge on the same goal without being interchangeable. Althea Labs covers both under a single visibility program.
Does llms.txt help with AI visibility? Not in Google Search. Google says the file “won’t negatively or positively impact your visibility or rankings,” and that Google Search does not use it. Other services may read it, which is why we keep one on our own site.
Is there an ideal page length for AI answers? Google says there is not. Match the length to the question and the evidence it needs, put the direct answer early, and stop when the answer is complete.
Does schema markup get a page cited? Not on its own. Google keeps structured data as part of ordinary SEO because it makes a page eligible for rich results, and Bing says accurate markup may support clearer grounding. Neither engine ties a citation to it.
How long before a structure change shows up? Nobody can give you a date for a heading rewrite on its own. Record a baseline in the first month of tracking, then compare the changed pages with an unchanged group over the same period. For the whole program, earned coverage included, we expect early wins in 60 to 90 days and compounding over months.
Start with an audit, not a rewrite
Your pages may already have clear headings and direct answers. Restructuring them will not fix weak source signals, blocked crawlers or thin earned media, and our position is that a rewrite bought before the diagnosis is money spent on the wrong layer.
Althea Labs is a brand presence and AI visibility agency that builds targeted visibility systems for B2B startups through integrated PR, SEO, and content programs. We call the method PR Plus. It runs in five steps, in order: Audit, ICP + Messaging, Visibility Tracking, Strategy, Execution. The audit comes first because it tells you whether the limit is your page content, your technical foundation, your PR footprint, or your community presence, and it gives you an entity-recognition baseline to measure against.
Find the missing layer before you approve the rewrite. Start your audit.
CONTINUE READING



