Royking Niba

Schema Markup for AI Search: What Google Actually Says You Need

· Royking Niba

Stock photograph of coloured lines of code displayed on a computer monitor. A generic illustration, not a screenshot of any specific site's markup.

There is a growing market in “AI schema”, “LLM-ready structured data” and audits that will sell you a new markup layer for AI search. Before you buy any of it, read the sentence Google publishes in its own documentation on AI features in Search: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.”

That is as direct as the documentation gets on any subject. It does not mean structured data is useless. It means the thing being sold as the requirement is not a requirement. This piece separates the two.

What the documentation actually states

Three statements from Google’s AI features documentation carry the whole argument:

  • “The best practices for SEO remain relevant for AI features in Google Search (such as AI Overviews and AI Mode).”
  • “There are no additional requirements to appear in AI Overviews or AI Mode.”
  • “You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.”

The controls Google names are the ones that already existed: nosnippet, data-nosnippet, max-snippet and noindex. There is no new opt-in tag, and the opt-out mechanism is the preview control set you have had for years.

Claim versus documentation

These are the five claims I have been asked to implement by clients in the last few months, each set against what the published guidance says.

The claimWhat the documentation saysVerdict
You need new AI-specific schema types to be cited in AI Overviews“There’s also no special schema.org structured data that you need to add”Contradicted outright
An llms.txt file is required for AI visibility in Google“You don’t need to create new machine readable files, AI text files, or markup”Contradicted outright
Structured data has replaced content quality as the ranking input“The best practices for SEO remain relevant for AI features”Unsupported. The stated position is continuity, not replacement
Marking up claims that are not on the page helps machines understand it“Don’t mark up content that is not visible to readers of the page”A named guideline violation
A structured data manual action will tank your rankings“A structured data manual action means that a page loses eligibility for appearance as a rich result; it doesn’t affect how the page ranks in Google web search”Overstated. The penalty is real but narrower than claimed

So what is schema markup still doing

Three jobs, all of them real, none of them the one being marketed.

1. Rich result eligibility

This is the documented, mechanical payoff and it has not changed. Valid markup in a supported type makes a page eligible for a rich result. Eligible is the operative word: Google states the markup earns eligibility, not the appearance itself. Formats are limited to three, with Google’s stated preference clear: “mark up your site’s pages using one of three supported formats: JSON-LD (recommended), Microdata, RDFa.”

2. Entity disambiguation

This is where structured data quietly earns its keep in an AI-heavy results page, and it is not about the AI layer at all. Markup states, unambiguously, which organisation published a page, which named person wrote it, and what that person is elsewhere on the web. Prose says the same thing, but prose has to be parsed and can be misread. A sameAs array cannot.

3. Making the facts extractable

Any system summarising a page has to decide what the page asserts. Markup that mirrors what is visibly on the page gives a second, structured statement of the same facts. It does not create an entitlement to be cited. It removes one way of being misread, which is a smaller claim and a true one.

A worked example: the markup I actually ship

For an expert-authored article, this is the whole of it. Article, a named Person with external profiles, and the publishing Organization. Nothing invented, every value also visible on the rendered page.

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Schema Markup for AI Search: What Google Actually Says You Need",
  "datePublished": "2026-09-25",
  "dateModified": "2026-09-25",
  "author": {
    "@type": "Person",
    "name": "Royking Niba",
    "jobTitle": "SEO and GEO consultant",
    "url": "https://roykingniba.com/about/",
    "sameAs": ["https://www.linkedin.com/in/roykingniba/"]
  },
  "publisher": {
    "@type": "Organization",
    "name": "Royking Niba",
    "url": "https://roykingniba.com/"
  },
  "mainEntityOfPage": "https://roykingniba.com/schema-markup-for-ai-search/"
}
</script>

What is deliberately absent is as important as what is there. No FAQ block for questions the page does not answer in visible text. No invented ratings. No Organization fields claiming awards or affiliations that appear nowhere on the site. Each of those would be a direct violation of the guideline Google states as “Don’t mark up irrelevant or misleading content, such as fake reviews or content unrelated to the focus of a page.”

The guidelines that bite

Google’s structured data general guidelines are short and the enforceable parts are worth knowing verbatim:

  • “Don’t block your structured data pages to Googlebot using robots.txt, noindex, or any other access control methods.” Markup on a page that cannot be fetched does nothing.
  • “Don’t mark up content that is not visible to readers of the page.” This is the single most common violation I find, usually through a plugin generating fields from a template.
  • “Don’t use structured data to deceive or mislead users. Don’t impersonate any person or organization, or misrepresent your ownership, affiliation, or primary purpose.”
  • “Provide up-to-date information. We won’t show a rich result for time-sensitive content that is no longer relevant.”
  • “Provide original content that you or your users have generated.”

The enforcement is narrow and specific. A structured data manual action costs the page its rich result eligibility and, per Google, “doesn’t affect how the page ranks in Google web search.” It shows in the Manual Actions report in Search Console like any other. Narrow does not mean harmless: on a site whose click-through rate depends on review stars or product pricing in the result, losing eligibility is a revenue event.

The checklist worth running instead

  1. Confirm every marked-up value also appears in the visible rendered page. Fix the mismatches before adding anything new.
  2. Confirm the marked-up page is fetchable by Googlebot and not blocked or noindexed.
  3. Use JSON-LD, and use one implementation. Two plugins each emitting an Article block is a conflict, not a reinforcement.
  4. Name a real author with a real profile page, and use sameAs to point at profiles that genuinely belong to them.
  5. Keep dateModified honest. Bumping it without changing the content is the kind of freshness signal that stops meaning anything once it is automated.
  6. Check the Manual Actions report. It is free, it takes thirty seconds, and it is the only place a structured data penalty is stated outright.
  7. Then go back to the content, because the documentation says the best practices for AI features are the ones you already had.

Related reading

Sources: Google Search Central documentation on AI features in Google Search and the structured data general guidelines, both checked 25 September 2026. Quoted sentences are Google’s own wording.

Leave a Reply

Your email address will not be published. Required fields are marked *