KEY TAKEAWAYS
- → +2.2% Change in ChatGPT citations after 1,885 pages added JSON-LD, statistically indistinguishable from zero. (Ahrefs, 2026)
- → 97% Share of valid llms.txt files that received zero requests in May 2026, across 137,000 domains. (Ahrefs, 2026)
- → 569 million Monthly GPTBot requests on Vercel’s network, with no evidence of JavaScript being executed by any major AI crawler. (Vercel and MERJ, 2024)
- → 25.7% How much fresher AI-cited URLs are than URLs in organic Google results, across 17 million citations. (Ahrefs, 2025)
- → 5 million AI-cited URLs in Semrush’s technical SEO study, which reports correlations and says outright they do not prove causation. (Semrush, 2025)
- → 10.13% Share of 300,000 domains with an llms.txt file; the file showed no correlation with AI citation frequency. (SE Ranking, 2025)
Most advice on technical SEO for AI citations mixes three very different kinds of evidence and presents them with the same confidence. A controlled test, a correlation across millions of URLs and a number nobody can trace all get quoted in the same breath, often in the same carousel. So I went back to the primary studies and sorted every popular claim by what actually stands behind it.
The short version: the only technical factor with strong evidence of a direct effect is whether AI crawlers can read your HTML at all. Schema markup, the favourite fix of the last two years, has now been tested head-on and did not move citations for pages already being cited. 53% of AI-cited pages carry schema markup, yet 1,885 pages that added it saw no statistically significant citation gain in ChatGPT or Google AI Mode. (Ahrefs, 2026). That gap between what cited pages have and what causes citation runs through every section below.
Each section below answers one question a learner or client is likely to ask, then shows the numbers behind the answer and where they came from. I have separated three grades of evidence throughout: controlled tests that compare changed pages with unchanged ones, large correlation studies that describe what cited pages have in common, and figures that circulate widely with no study behind them. Only the controlled tests can tell you what to change.
1 Does Adding Schema Markup Increase AI Citations? The 1,885-Page Test
No, not for pages AI already cites. Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 matched controls. ChatGPT citations moved +2.2% and AI Mode +2.4%, both indistinguishable from zero, while AI Overviews fell -4.6%. (Ahrefs, 2026)
| METRIC | VALUE | SOURCE |
|---|---|---|
| Pages that added JSON-LD | 1,885 | Ahrefs |
| Matched control pages | 4,000 | Ahrefs |
| Change in Google AI Overviews citations | -4.6% | Ahrefs |
| Change in Google AI Mode citations | +2.4% | Ahrefs |
| Change in ChatGPT citations | +2.2% | Ahrefs |
This is the most important study in the whole topic because it is the only one built to separate cause from coincidence. Each treated page was matched with control pages from other domains that had similar citation levels and never added schema. If schema worked, the treated group should have pulled ahead. It did not.
The AI Overviews dip deserves context rather than alarm. Ahrefs puts the loss at around 12 daily citations per page, in a sample where most pages were getting hundreds, and both groups were already falling before anything changed. The honest reading is that schema had no clear effect in either direction.
There is one real limit. Every page in the study was already inside AI’s consideration set. The study cannot say whether schema helps a page nobody cites yet get picked up, and it only looked at schema in the HTML, not schema injected by JavaScript.
If you want certainty for your own site, run a small version of the same test. Pick a handful of pages that already earn some AI citations, add schema to half of them, leave the other half alone, and compare both groups after a month. If they move together, the platform moved, not your markup.
2 Why Do 53% of AI-Cited Pages Have Schema If It Doesn’t Cause Citations?
Because the sites that add schema also do everything else well. Ahrefs found AI-cited pages are Almost 3X as likely to carry JSON-LD, and Semrush found Organization schema on 34% of pages cited in AI Mode. Both are correlations, and both studies say so. (Semrush, 2025)
| METRIC | VALUE | SOURCE |
|---|---|---|
| AI-cited URLs in Semrush’s study | 5 million | Semrush |
| Organization schema, AI Mode citations | 34% | Semrush |
| Organization schema, ChatGPT citations | 25% | Semrush |
| FAQ schema, AI Mode citations | 5.5% | Semrush |
| Citations for the best slug-length band | 87,000 | Semrush |
| Slug range that performed best | 17-40 characters | Semrush |
Semrush’s study is useful, and I use Semrush daily, but its own authors are careful: the patterns in 5 million cited URLs point to strong technical foundations, not to any single lever. Well-maintained sites add schema, write tidy URLs, load quickly, earn links and publish better content. AI systems cite those sites. Remove schema from the picture and the rest of those signals still carry the page.
The slug-length finding shows why correlations need care. URLs with 17-40 characters slugs collected the most citations, but short slugs are often homepages and very long ones are often deeply nested or stuffed. Rewriting your URLs to hit a character count would cost you redirects and gain you nothing a study has shown.
The takeaway for learners is simple. When a statistic describes what cited pages have, read it as a description of good sites, not a recipe. Keep schema for rich results and entity clarity. Just stop selling it to clients as an AI citation fix.
3 Can AI Crawlers Read JavaScript? What 569 Million GPTBot Requests Showed
Mostly not. Vercel and MERJ measured 569 million GPTBot requests in one month and found none of the major AI crawlers render JavaScript. ChatGPT’s crawler downloaded JavaScript files in 11.50% of fetches and Claude’s in 23.84%, without running them. Gemini is the exception. (Vercel and MERJ, 2024)
| METRIC | VALUE | SOURCE |
|---|---|---|
| GPTBot requests in one month | 569 million | Vercel and MERJ |
| Claude crawler requests in one month | 370 million | Vercel and MERJ |
| ChatGPT fetches that were JavaScript files | 11.50% | Vercel and MERJ |
| Claude fetches that were JavaScript files | 23.84% | Vercel and MERJ |
| ChatGPT fetches hitting not-found pages | 34.82% | Vercel and MERJ |
| Claude fetches hitting not-found pages | 34.16% | Vercel and MERJ |
| Googlebot fetches hitting not-found pages | 8.22% | Vercel and MERJ |
This is the one technical factor where the evidence and the logic line up. If your main content only appears after client-side rendering, GPTBot, ClaudeBot and PerplexityBot receive an empty shell. No amount of content quality fixes a page the crawler never sees. Server-side rendering, static generation or pre-rendering for key pages is the fix.
Two cautions on dates and wording. This study was published in December 2024, so crawler behaviour may have shifted since. And the finding has drifted as it spread. The original says Vercel saw 569 million GPTBot requests and that no major AI crawler renders JavaScript. Many articles now say a study of five hundred million fetches found zero JavaScript execution, which is close but not what was written.
You can check your own site without trusting anyone. Your server logs show which AI bots fetch which URLs and which responses they get. My step-by-step log file analysis process works for AI crawlers too; just filter by their user agents and compare what they fetched against what a browser renders.
4 Does llms.txt Get You Cited? 97% of Files Received Zero Requests
Not on current evidence. Ahrefs checked 137,000 domains and found 97% of valid llms.txt files got no requests at all in May 2026. SE Ranking studied 300,000 domains and found no link between having the file and citation frequency. (Ahrefs, 2026)
| METRIC | VALUE | SOURCE |
|---|---|---|
| Domains in SE Ranking’s study | 300,000 | SE Ranking |
| Domains with an llms.txt file | 10.13% | SE Ranking |
| Adoption among high-traffic sites | 8.27% | SE Ranking |
| Domains in Ahrefs’ server-log study | 137,000 | Ahrefs |
| Domains publishing llms.txt (Ahrefs sample) | 28% | Ahrefs |
| Valid files with zero requests in May 2026 | 97% | Ahrefs |
I deployed llms.txt on this site myself, so I have no reason to talk it down. But two very different methods now agree. SE Ranking correlated adoption with citations and found nothing, and removing the file from its model actually improved the model’s accuracy. Ahrefs read server logs and found the bots barely ask for the file.
Adoption figures vary wildly, from 10.13% in SE Ranking’s sample to 28% among Ahrefs’ more technical customers. Neither describes the whole web, so treat any single adoption number with care.
What the logs do show is worth acting on. AI bots fetch your normal HTML pages constantly, so the effort that would go into polishing an llms.txt file is better spent making sure those pages load cleanly, return the right status codes and say something specific enough to quote.
My position: publishing llms.txt is cheap housekeeping and harmless. Just do not put it in a proposal as a citation lever, and do not let it replace the work that the rendering and content sections of this page point to.
5 Does Content Freshness Affect AI Citations? 17 Million Citations Compared
There is a measurable preference, though it is a correlation. Across 17 million citations, AI-cited URLs averaged 1064 days old against 1432 days for organic Google results, making them 25.7% fresher. ChatGPT’s citations were 458 days newer than organic results. (Ahrefs, 2025)
| METRIC | VALUE | SOURCE |
|---|---|---|
| Citations analysed across platforms | 17 million | Ahrefs |
| Average age of AI-cited URLs | 1064 days | Ahrefs |
| Average age of organic Google URLs | 1432 days | Ahrefs |
| How much fresher AI-cited URLs are | 25.7% | Ahrefs |
| How much newer ChatGPT citations are than organic | 458 days | Ahrefs |
Freshness is where good practice and hype collide. The data supports keeping important pages genuinely current, and ChatGPT shows the strongest pull towards newer pages. Google AI Overviews behaved differently, citing pages slightly older than the organic results.
What the data does not support is changing a date without changing the page. The study measured publication and update dates, not whether an edit caused a citation. Update when the facts change, add what is new, and record what you changed. That is the same discipline I recommend for anyone trying to write content that ranks on Google and gets cited by AI.
Keep the averages in proportion too. AI-cited pages were still close to three years old on average, so long-lived reference content keeps its value.
6 Which Technical AI SEO Claims Have No Primary Source?
Several popular numbers have no study you can open. A 47% (commonly cited; no primary source traceable) citation lift for structured data, a claim that 69% (commonly cited; no primary source traceable) of AI crawlers cannot run JavaScript, and a 3× (commonly cited; no primary source traceable) citation loss for pages not updated quarterly all circulate without a traceable sample or method. (Snezzi (repeating an untraced figure), 2026)
| CLAIM | VALUE | SOURCE |
|---|---|---|
| Citation lift from clear structured data | 47% (commonly cited; no primary source traceable) | Snezzi (repeating an untraced figure) |
| AI crawlers unable to execute JavaScript | 69% (commonly cited; no primary source traceable) | SwingIntel (attributed to a 2025 searchVIU study) |
| Citation loss for pages not updated quarterly | 3× (commonly cited; no primary source traceable) | Dotcom-Monitor |
None of these may be wrong in spirit. JavaScript rendering really is a problem, and stale pages really can fade. The issue is precision. A specific percentage suggests someone measured it, and when a learner repeats it in an audit or a client pitch, they are vouching for a study that may not exist.
Page speed is the clearest blind spot. It is widely listed as an AI citation factor, and speed genuinely helps users and Google rankings, which feed AI Overviews indirectly. But I could not find a public controlled test isolating speed or Core Web Vitals as a cause of AI citations. Until one exists, call it a supporting factor, not a proven lever.
Here is a simple test before you quote any AI SEO number: can you name the organisation, the sample size and the year, and open the page where it lives? If not, leave it out or label it. That habit is the backbone of the approach in my Hybrid Engine Optimisation guide, and it will make your recommendations harder to argue with.
Search every figure in this article
Methodology
I reviewed the primary studies behind the most repeated technical SEO claims for AI citations and opened every source cited here on the research date. Facts are graded as verified when the figure appears on the source page and unverified when it circulates without a traceable study. Several research stages ran in a reduced mode: SERP checks ran through web search, Source verification was done by opening each page manually and the charts were drawn locally from the data.
- Sources consulted: 41
- Sources cited: 9
- Data freshness: current year: 4, last year: 3, older: 1
- Data range: 2024-12-17 to 2026-09-29
- Research date: 2026-09-29
- Update schedule: Quarterly
- Limitations: Every study here measures pages already visible to AI systems or correlates traits of cited pages. No public controlled test isolates page speed, Core Web Vitals or robots.txt changes as causes of AI citations.
Frequently Asked Questions
Does schema markup help you get cited by ChatGPT or AI Overviews?
Not measurably, for pages already being cited. Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 controls and found ChatGPT citations moved +2.2% and AI Mode +2.4%, both indistinguishable from zero. Schema still earns rich results, so keep it for that reason. (Ahrefs, 2026)
Do AI crawlers render JavaScript?
The major ones tested do not. Vercel and MERJ found GPTBot, ClaudeBot and PerplexityBot download JavaScript files (11.50% of ChatGPT’s fetches) without executing them. Gemini is the exception because it uses Googlebot’s rendering. If your main content only appears after client-side rendering, most AI crawlers see an empty page. (Vercel and MERJ, 2024)
Is llms.txt worth adding for AI visibility?
Not for citations, based on current data. Ahrefs found 97% of valid llms.txt files got zero requests in May 2026, and SE Ranking’s study of 300,000 domains found no link between the file and citation frequency. It is cheap to publish, so treat it as optional housekeeping rather than a ranking lever. (Ahrefs, 2026)
Does updating content help AI citations?
There is a freshness preference, though it is a correlation. Across 17 million citations, Ahrefs found AI-cited URLs were 25.7% fresher than organic results, with ChatGPT favouring the newest pages. Google AI Overviews cited slightly older pages than the organic results did. Update pages when the content genuinely changes, not just the date. (Ahrefs, 2025)
Does page speed affect AI citations?
No controlled study found for this article isolates page speed as a cause of AI citations. Claims such as pages losing citations at 3× the rate without quarterly updates circulate without a named study or sample. Speed still matters for users and Google rankings, which feed AI Overviews indirectly. (Dotcom-Monitor, 2026)
Sources & References
- Ahrefs. “We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved..” ahrefs.com/blog/schema-ai-citations/. Accessed 2026-09-29.
- Semrush. “How Do Technical SEO Factors Impact AI Search? [Study].” semrush.com/blog/technical-seo-impact-on-ai-search-study/. Accessed 2026-09-29.
- Vercel and MERJ. “The rise of the AI crawler.” vercel.com/blog/the-rise-of-the-ai-crawler. Accessed 2026-09-29.
- SE Ranking. “Does LLMs.txt impact your AI visibility and citations? No, according to research.” seranking.com/blog/llms-txt/. Accessed 2026-09-29.
- Ahrefs. “What Is llms.txt, and Should You Care About It?.” ahrefs.com/blog/what-is-llms-txt/. Accessed 2026-09-29.
- Ahrefs. “AI Assistants Prefer to Cite Fresher Content (17 Million Citations Analyzed).” ahrefs.com/blog/do-ai-assistants-prefer-to-cite-fresh-content/. Accessed 2026-09-29.
- Snezzi. “How AI Chatbots Pick Sources (and How to Get Cited).” snezzi.com/blog/how-ai-chatbots-pick-sources-an-inside-look-marketers/. Accessed 2026-09-29.
- SwingIntel. “Technical SEO for AI Search: Schema, Rendering and Citation Factors (2026).” swingintel.com/blog/technical-seo-factors-ai-search. Accessed 2026-09-29.
- Dotcom-Monitor. “How Website Speed Impacts SEO and AI Search in 2026.” dotcom-monitor.com/blog/website-speed-affect-seo/. Accessed 2026-09-29.
Last updated:
Comments are closed