Skip to content
Field notes
SEOMarketing

Structured Data in 2026: What Schema Markup Still Does for SEO and AI Search

Google stopped rendering FAQ rich results in May 2026, and a controlled Ahrefs study of 1,885 pages found that adding JSON-LD moved AI citations by less than 5% on every platform tested. Structured data is not dead, but its job has changed. Here is what schema markup still buys you, which types Google retired, why AI assistants ignore JSON-LD during live retrieval, and how much a marketing team should actually invest.

August 4, 202619 min
Structured data in 2026 concept showing a JSON-LD code block on a web page connecting to rich results, AI answers and a knowledge graph entity cluster, with one connection faded
ShareShare on XLinkedIn

Structured data has spent twenty years as the safest recommendation in SEO. It was cheap to add, it never hurt, and it occasionally produced a star rating or a set of expandable questions that made your listing look bigger than everyone else's. Nobody had to defend the line item.

2026 broke that comfort. Google stopped rendering FAQ rich results on May 7. Practice Problem support disappeared in January. Seven schema types were retired the year before. And in May, Ahrefs published a controlled study of 1,885 pages that added JSON-LD and found that AI citations barely moved on any platform. If your structured data strategy was built on the assumption that more markup equals more visibility, the assumption is now measurable, and it did not hold up. What follows is what structured data still does, what it stopped doing, and how to spend the right amount of effort on it.

Table of Contents

What Structured Data Actually Is

Structured data is a machine-readable description of a page's contents, written in a shared vocabulary from schema.org and embedded so that a crawler does not have to infer meaning from layout. A recipe page tells a human it is a recipe through a photo and an ingredient list. Structured data tells a machine the same thing through explicit fields: cook time, yield, calories, author.

Three formats exist. JSON-LD sits in a script tag and is the format Google recommends and nearly everyone uses. Microdata and RDFa are inline attributes wrapped around visible HTML, older and increasingly rare. When people say schema markup in 2026, they almost always mean JSON-LD.

The important distinction, and the one most teams get wrong, is between the vocabulary and the search feature. Schema.org defines hundreds of types. Google supports rich results for a much shorter list, and that list shrinks and grows on Google's schedule, not schema.org's. FAQPage is still a perfectly valid schema.org type today. It just no longer produces anything visible in Google Search. Valid markup and rendered markup are two different things, and confusing them is how teams end up defending budget for work that stopped producing output eighteen months ago.

The 2026 Deprecation List

The clearest signal of where structured data is heading is what Google has stopped rendering. The pattern is consistent: features that were widely abused, rarely clicked, or duplicative of what the AI layer now does have been quietly retired.

Schema typeStatusWhen
How-toRich result removed2023
Seven lesser-used types (including Sitelinks Search Box, Book Actions, Course Info and others)Support retiredJune 2025
Practice ProblemReporting removedJanuary 2026
FAQPageRich result removed, type still validMay 7, 2026
Organization, Product, Offer, Article, BreadcrumbList, ReviewFully supportedCurrent

The FAQ removal deserves the most attention because it is the one most content teams built process around. Google updated its documentation on May 7, 2026 with a deprecation notice and no accompanying blog post or explanation. FAQ rich results stopped appearing that day. The Search Console search appearance filter, the rich result report, and Rich Results Test support were scheduled to follow in June, and teams pulling FAQ data through the Search Console API were given until August to adjust their calls. That deadline lands this month, which is a good reason to check whether any of your reporting pipelines still depend on it.

None of this affects rankings. Google has been explicit that removing a rich result changes what the SERP looks like, not how pages are ordered. The loss is click-through rate on listings that used to occupy more vertical space, and that loss is real but bounded.

It is also worth resisting the overcorrection. When Google announced the January 2026 removals, a widely shared Reddit thread asked whether Google was killing schema.org entirely. John Mueller's answer was measured: markup types come and go, and a precious few are worth holding on to. That is the accurate reading. Google is pruning features, not abandoning structured data. If you are auditing your own coverage as part of a broader content audit, treat deprecated types as cleanup, not as evidence that the whole practice is over.

The Ahrefs Study: What the Data Showed

For two years the case for schema as an AI visibility lever rested on a single correlation, and it was a striking one. Ahrefs analyzed six million URLs and found that pages cited by AI were almost three times more likely to carry JSON-LD than pages that were not. That statistic travelled well. It appeared in conference decks and LinkedIn carousels as proof that markup drives citations.

Ahrefs then did the thing that correlations deserve and tested it properly. The team pulled crawler history to identify 1,885 pages that transitioned from no JSON-LD to having JSON-LD between August 2025 and March 2026. For each treated page they selected three control pages from different domains with similar pre-period citation levels that never added schema. Then they measured citations across Google AI Overviews, AI Mode, and ChatGPT for 30 days before and 30 days after the change, using a matched difference-in-differences design to strip out platform-wide trends.

PlatformEffect on citationsVerdict
Google AI Overviews-4.6%Small but statistically significant decline versus matched controls
Google AI Mode+2.4%Indistinguishable from zero
ChatGPT+2.2%Indistinguishable from zero

The AI Mode number is the most instructive. Raw before-and-after growth for treated pages came in at +43%, which would have made a spectacular headline. Control pages gained almost as much. AI Mode was expanding for everyone that quarter. Strip the platform trend out and the +43% collapses to +2.4%, which is noise. This is precisely the trap that most agency case studies fall into, and it is why incrementality testing matters as much in SEO as it does in paid media.

Ahrefs ran four separate tests, including an event study to confirm the two groups were not already diverging before treatment and a re-run with a symmetrical window excluding the recrawl period. All four agreed. The most consistent finding was that not much changed.

The caveat matters as much as the headline. Every page in the dataset already had 100+ AI Overview citations before schema was added. These were pages the models were already retrieving. The study cannot say anything about pages that are invisible to AI systems entirely, where markup might still help with crawling, parsing, or indexing. If you are working on AI visibility from a standing start, the study does not tell you schema is useless. It tells you schema will not push a page that is already being cited any higher.

Ahrefs was also careful about what it did not isolate. Pages that add JSON-LD often change other things at the same time. All schema types were pooled together, so it remains possible that Organization behaves differently from Article. The measurement window was 30 days, which may miss slow-burn effects. And only schema present in the raw HTML was tested, not schema injected by JavaScript, which AI crawlers appear to treat differently.

What AI Systems Read When They Fetch Your Page

The Ahrefs result becomes much less surprising once you look at a separate experiment from searchVIU, a German technical SEO agency. They built otherwise-identical pages that differed only in the presence of schema, in all three formats, and asked ChatGPT, Claude, Gemini, Perplexity, and Google AI Mode to extract specific data from them.

In the cleanest version of the test, a product price was placed exclusively inside JSON-LD and nowhere in the visible page. Zero of the five systems retrieved it. Across the board, during live retrieval, every system extracted only visible HTML. JSON-LD, hidden Microdata, and hidden RDFa were all ignored.

This does not mean schema is invisible to AI. searchVIU was explicit that markup may still be consumed earlier in the pipeline, during training, indexing, or search-index lookup, and Google's crawlers certainly extract it. The distinction is between retrieval-time reading, where schema is ignored, and index-time understanding, where it still contributes. That distinction explains why the citation numbers barely moved while the correlation between schema and citation remains high: both are downstream of the same underlying quality signals.

It also reframes what actually drives citations. Ahrefs' own research on the same platforms found freshness and content quality to be far stronger predictors. If you want to be cited, the levers are the ones covered in answer engine optimization and GEO versus SEO: extractable answers in visible prose, clear headings, current data, and demonstrable authority on the topic. Markup is not on that list. Neither, for what it is worth, is llms.txt, which has attracted similar unearned enthusiasm.

Where Structured Data Still Earns Its Keep

None of this makes schema a waste of time. It makes it a plumbing investment rather than a growth lever, and those are budgeted differently. Here is the honest inventory of what markup still buys you in 2026.

Rich results that survived

Product, Offer, Review, and AggregateRating still produce price, rating, and review count in the SERP. For ecommerce this is a straightforward CTR play with no ambiguity about whether it renders. BreadcrumbList still replaces the raw URL string with a readable path. Article and BlogPosting still feed author, headline, and publish date into Top Stories and Discover surfaces. Video, Event, Recipe, and Job Posting remain intact in their verticals.

Entity resolution and the knowledge graph

Organization schema is the most undervalued type on the list. It states your legal name, logo, address, contact points, social profiles, and relationships in a form that requires no inference. That feeds the knowledge panel, and more importantly it feeds every downstream system that needs to decide whether the MarqOps in one document is the same MarqOps in another. This is the core of entity SEO, and it is the one area where the machine-readable format genuinely does work that visible prose cannot do as reliably.

Disambiguation at scale

The larger and more repetitive your site, the more schema pays. A 40-page brochure site gains little. A site with 12,000 product pages, or a publisher with fifteen years of archives, needs a consistent way to tell systems which pages are the same type of thing, how they relate, and which entity they belong to. Teams building topical authority across large clusters get real value here, because markup makes the cluster structure explicit rather than implied.

Schema typeWhat it still buysPriority
OrganizationKnowledge panel, entity resolution, brand disambiguationHigh for every site
Product + OfferPrice and availability in SERP, Merchant Center eligibility, Shopping GraphHigh for ecommerce
Article / BlogPostingAuthor, date, and headline metadata for Discover and Top StoriesMedium for publishers
BreadcrumbListReadable navigation path in the SERPMedium, cheap to add
Review / AggregateRatingStar ratings where still eligibleMedium, abuse-sensitive
FAQPageNothing visible in Google Search since May 2026Low, safe to leave in place
How-toNothing since 2023Remove

Product Data, Merchant Listings, and Agentic Checkout

Ecommerce is the exception to almost everything above, because in commerce the structured data is not decoration on a web page. It is the product itself, expressed as data.

Google's Shopping Graph reportedly holds around 60 billion product listings, and Universal Cart launched at I/O 2026 across the US, Canada, and Australia. When an AI agent completes a purchase on a shopper's behalf, it is not reading your product page prose. It is consuming structured product data, and the quality of that data determines whether your item is a candidate at all. Google has also added properties to its Merchant Listing guidance, including Product.category for classifying items directly in on-page markup, and now recommends the OnlineStore subtype of Organization rather than the generic type for retail sites.

For retailers this changes the calculus completely. Product and Offer markup is not an SEO nicety, it is the interface to a purchasing surface. The same logic that governs feed quality in Google Shopping ads now governs organic and agentic visibility: complete attributes, valid identifiers, accurate price and availability, and no drift between what your feed says and what your page says. A mismatch that gets a product disapproved in Merchant Center is the same mismatch that makes an agent skip you.

Structured data in 2026 infographic showing the Ahrefs schema study results across AI Overviews, AI Mode and ChatGPT, the schema types Google stopped rendering, the types that still earn rich results, and where schema works in the crawl and indexing pipeline

Structured data in 2026: what the controlled data showed, what Google retired, and what still pays.

Implementing Schema Without Burning a Sprint

The single biggest cost mistake in structured data is treating it as a page-level task. Hand-writing JSON-LD per page, or worse, asking writers to paste it into a CMS field, guarantees drift, breakage, and a maintenance burden that grows with the site. Structured data belongs at the template level.

Generate it from the data you already have

Your CMS already knows the author, publish date, headline, and category of every article. Your commerce platform already knows the price, availability, GTIN, and brand of every product. Schema should be assembled from those fields at render time, so that changing a price updates the markup automatically. If a human has to remember to update JSON-LD, it will be wrong within a quarter.

Keep it in the HTML, not the client bundle

Schema injected by client-side JavaScript is treated differently by AI crawlers than schema present in the server-rendered HTML, and there is no upside to the harder path. Server-render it.

Mark up the primary purpose of the page, nothing else

Google's March 2026 changes targeted exactly this: markup describing content that was not the main point of the page. A product page marked up as a Product is correct. A blog post marked up with Product schema because it mentions a product is the abuse pattern that got several rich result types killed in the first place.

Mirror everything in visible text

Given what searchVIU found, treat schema as a second copy of facts that already exist on the page rather than the only copy. This costs nothing when your markup is generated from page data, because the same fields drive both.

Automate validation, not authoring

Run the Rich Results Test and Schema Markup Validator in CI against representative page templates, not against every URL. One broken template is worth catching. Ten thousand individually validated pages is a report nobody reads. This is a natural fit for the kind of continuous checking covered in SEO automation, where the value is in catching regressions early rather than in producing more output.

A 30-Day Structured Data Audit

If you have not looked at your markup since 2024, there is almost certainly dead weight in it. Here is a sequence that fits inside a month without derailing anything else.

Week 1: inventory and triage. Crawl the site and extract every JSON-LD block by type and template. You are looking for three things: types Google no longer renders, types applied to the wrong page purpose, and templates where markup is missing entirely. Check whether any Search Console API calls in your reporting still request the FAQ appearance filter, because that pipeline breaks this month.

Week 2: fix the load-bearing types. Get Organization right first, including logo, sameAs profiles, contact points, and the OnlineStore subtype if you sell. Then Product and Offer if you are in commerce, then Article or BlogPosting, then BreadcrumbList. These four cover most of the remaining value.

Week 3: remove and consolidate. Strip How-to markup. Leave FAQPage in place if it is generated automatically and costs nothing, since it is still valid and harmless, but stop writing FAQ content specifically to trigger a rich result that no longer exists. Keep writing FAQ sections for readers and for extractability in AI Overviews, which pull from visible text, not from markup.

Week 4: wire up monitoring. Add template-level validation to your deploy pipeline and set an alert on Search Console structured data errors. Then, critically, set a baseline for the metrics you will use to judge whether any of this mattered.

Mistakes That Waste Schema Effort

Treating schema as an AI citation strategy. This is now the expensive one. The controlled data says it does not work that way, and every hour spent expanding markup coverage in pursuit of citations is an hour not spent on the freshness and content quality that demonstrably do drive them.

Marking up content that is not on the page. Google's guidelines are unambiguous, and the enforcement pattern of the last three years has been to remove features rather than penalize sites. The cost of abuse is that everyone loses the feature.

Running three plugins that each emit their own JSON-LD. Duplicate and conflicting Organization blocks are common on WordPress sites that have accumulated SEO plugins over the years. Pick one source of truth.

Chasing every new schema.org type. Schema.org has hundreds of types Google has never supported and never will. Support the ones with a documented search feature or a documented consumer, and ignore the rest.

Letting price and availability drift. For retailers this is the most expensive error on the list, because it does not just cost a rich result. It costs Merchant Center approval, Shopping eligibility, and agent consideration in one move.

Confusing valid with useful. A green checkmark in the Rich Results Test means your syntax parses. It says nothing about whether the feature still exists or whether anyone sees it.

How to Tell Whether Schema Is Doing Anything

The honest answer for most sites is that you cannot tell from aggregate data, and that is exactly why the industry believed a correlation for two years. If you want a real answer for your own site, run the small version of what Ahrefs ran.

Pick five to ten pages where you plan to add markup, ideally ones already getting some visibility so you have a baseline. Pick a matched set of five to ten control pages with similar visibility that you will not touch. Record baseline impressions, clicks, and AI citations for both groups. Add schema to the test set only, note the date, and change nothing else on those pages during the window. Compare both groups after 30 days, and ask whether the treated pages moved more than the controls did. If both moved together, you measured a platform trend, not your change.

That design is more work than reading a rank tracker, which is why almost nobody does it. It is also the only way to distinguish a real effect from a seasonal one. Tracking AI citations alongside traditional rankings is getting easier as AI search visibility tools mature, and a unified view of impressions, citations, and conversions in one dashboard makes the comparison a five-minute exercise rather than a spreadsheet project. That consolidation is the reason MarqOps folds SEO, analytics, ads, and creative into a single system rather than leaving teams to reconcile seven tools by hand.

Frequently Asked Questions

What is structured data in SEO?

Structured data is a standardized, machine-readable description of a page's contents using the schema.org vocabulary, usually embedded as JSON-LD in a script tag. It tells search engines explicitly what a page is about, who published it, and what entities it references, rather than requiring them to infer that from layout and prose. It makes pages eligible for rich results and helps systems resolve your brand as a distinct entity.

Does schema markup improve rankings?

No. Google has stated consistently for over a decade that structured data is not a ranking factor. It affects eligibility for search features and how clearly your entities are understood, not the order of results. Any ranking improvement observed after adding schema is usually attributable to the other technical and content work that happened alongside it.

Does schema markup help you get cited by AI?

The best available evidence says no, at least for pages already visible to AI systems. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against 4,000 matched controls and found effects of -4.6% on AI Overviews, +2.4% on AI Mode, and +2.2% on ChatGPT. Only the first was statistically significant and it was a decline. Freshness, content quality, and authority remain the stronger levers.

Should I remove FAQ schema now that Google dropped FAQ rich results?

You do not need to. FAQPage remains a valid schema.org type and leaving the markup in place causes no harm, particularly if your CMS generates it automatically. What should change is the effort allocation: stop creating FAQ content specifically to trigger a rich result that stopped rendering on May 7, 2026, and check whether any Search Console API calls in your reporting still request the FAQ appearance filter, since that support ends in August.

Which schema types still matter in 2026?

Organization for entity identity and the knowledge panel, Product with Offer for commerce rich results and Merchant Center eligibility, Article or BlogPosting for author and date metadata, and BreadcrumbList for navigation paths in the SERP. Review and AggregateRating still render where eligible. Video, Event, Recipe, and Job Posting remain supported in their verticals. How-to and FAQ no longer produce anything visible in Google Search.

Do AI assistants read JSON-LD when they fetch a page?

Not during live retrieval. A searchVIU experiment placed a product price exclusively in JSON-LD and asked ChatGPT, Claude, Gemini, Perplexity, and Google AI Mode to extract it. None of the five did. All extracted only visible HTML. Schema may still be consumed earlier in the pipeline during crawling and indexing, but any fact you want an assistant to quote must appear in visible page text.

Is JSON-LD better than Microdata or RDFa?

For practical purposes yes. JSON-LD is Google's recommended format, it is separated from your presentation markup so it does not break when templates change, and it is what nearly every CMS and plugin emits. Microdata and RDFa remain valid but are harder to maintain and offer no advantage. Use JSON-LD server-rendered into the HTML rather than injected by client-side JavaScript.

How much should a marketing team invest in structured data?

Enough to get four types correct at the template level and keep them validated, and no more. For most sites that is a one-time engineering effort of a few days plus continuous monitoring. Ecommerce is the exception: product data quality is load-bearing for Shopping, Merchant Center, and agentic checkout, and deserves ongoing ownership rather than a one-off project.

The Bottom Line

Structured data spent a decade as the recommendation nobody had to justify, and 2026 turned it into one that requires a real argument. The argument still exists, it is just narrower than it was. Schema makes you eligible for the rich results that survived, it makes your brand resolvable as an entity, and in commerce it is the format that search engines and purchasing agents both consume. Those are genuine returns and they are worth the engineering time.

What it does not do is buy AI citations. The correlation that convinced the industry otherwise turned out to be exactly what it looked like: well-run sites do many things right at once, and markup was riding along rather than driving. When someone finally isolated the variable, the effect disappeared. That is a useful result, because it frees up attention for the things that do move the needle in AI search, and because the same discipline applies to every other tactic currently being sold as an AI visibility lever.

Get the four types right, automate them at the template level, validate on deploy, mirror every fact in visible text, and then go spend the reclaimed hours on content that is worth citing. That is the whole strategy, and it is smaller than it used to be, which is the point.

One useful operating system. Once a week.

Tactical notes across SEO, paid media, analytics, reporting, and AI. No filler.

Put the system to work

Turn marketing evidence into the next decision.

Explore the interactive sample, then start a self-service seven-day trial.