AI systems don’t see a web page the way a visitor does: they take the raw HTML the server sends, extract the text, split it into passages and match those passages to sub-questions derived from the buyer’s question. A page designed for people loses information in this process. Content loaded by JavaScript, text on images, information locked in PDFs, vague headings and paragraphs that don’t stand on their own are incomplete or missing in the machine’s view.
Picture a corporate homepage: a full-screen video with “Building the future together” on top, scrolling customer logos below, service details that open in tabs and a “Download catalog” button. An impressive experience for a visitor. For an AI crawler, a largely empty page.
The same page serves two different readers, and they don’t see the same thing. This article opens up what the machine sees and what it misses.
In what steps does AI read a page?
- Crawling. A crawler requests the page. OpenAI uses three separate crawlers: GPTBot for model training, OAI-SearchBot for the search index and ChatGPT-User for a user’s request inside a conversation. Each can be allowed or blocked separately in robots.txt.
- Raw HTML. The crawler takes the first HTML the server sends. According to Vercel and MERJ’s analysis of more than 500 million requests, GPTBot, ClaudeBot and PerplexityBot download JavaScript files but don’t execute them. Google is the exception: Gemini shares the rendering infrastructure of Googlebot.
- Text extraction. Menus, footers and repeated elements are stripped; the main text of the page remains. Images, videos and animations disappear at this stage unless they carry text.
- Passage splitting. The text is split into passages around headings and paragraphs. Each passage can be evaluated independently of the others.
- Matching. The buyer’s question is broken into sub-questions in the background; Google has described this as “query fan-out” in AI Mode. According to Ahrefs’ study of 1.4 million ChatGPT prompts, the strongest factor in whether a page is cited is the semantic similarity between its title and those sub-questions.
- Selection. In the same study, ChatGPT cited only about half of the pages it read. Being read doesn’t mean being used.
Seven gaps in pages written for people
1. Content loaded by JavaScript. Tabs, accordions, filtered product lists and pricing tables loaded later don’t exist for an AI crawler if they’re missing from the first HTML the server sends. Corporate sites built as single-page applications leave this gap most often.
2. Text on images. Customer logos, certificate images, infographics and slide screenshots are strong proof for people. For a machine, they’re often just image files. If “we work with more than 50 industrial companies” lives in the logos but isn’t written in the text, the model doesn’t know it.
3. Information locked in PDFs. Data sheets, catalogs and capacity tables are often offered as PDFs. Even when PDFs can be crawled, that content stays disconnected from the page, isn’t treated as part of it and is hard to keep current. Information critical to the decision should also sit in the HTML text of the page.
4. Creative but vague headings. “We add value”, “Solutions that make a difference”, “Our story”. These headings match no sub-question. “Our HACCP-compliant cold storage capacity for the food sector” matches directly.
5. Paragraphs that don’t stand on their own. The model reads the page in passages. “This solution meets the needs of the sector” says nothing without the previous paragraph. Each section should make sense on its own, naming the brand and the topic.
6. Missing numbers and conditions. Phrases like “large capacity”, “fast delivery” and “competitive pricing” don’t answer an elimination question. If the buyer asks for a company “that can produce 50 tons a month”, that number needs to be on the page. For pricing, stating at least the pricing model and the factors that affect it, in place of “contact us”, keeps the model from filling the gap with guesses.
7. Access blocks. Security plugins, CDN rules and robots.txt settings can block AI crawlers without anyone noticing. Since July 1, 2025, Cloudflare has blocked AI crawlers by default on new domains.
What the visitor sees vs. what the machine sees
| Page element | Visitor | AI crawler | Fix |
|---|---|---|---|
| Hero video and slogan | Strong first impression | One vague sentence | A clear paragraph defining the brand below it |
| Customer logos | Proof of trust | Image files | State sectors and customer count in text |
| Tabbed service details | Tidy content | Empty if loaded by JavaScript | Serve the content in the first HTML |
| Catalog PDF | Downloadable detail | A separate document detached from the page | Move critical information into page text |
| “Solutions that make a difference” heading | Brand voice | A heading that matches nothing | A heading that answers the question |
| “Call for pricing” | Sales handoff | Information gap | Pricing model and the factors behind it |
How do you test your site through a machine’s eyes?
- Turn off JavaScript. Disable JavaScript in the browser and open key pages. Every piece of information that disappears is lost to AI crawlers too.
- Read the page source. In “View page source”, do your critical facts, capacity, region, certification, appear as text?
- Read the headings on their own. List the page’s H2 headings one under another. Could a buyer tell from this list which questions the page answers?
- Copy one section and read it without context. Does it make sense alone, or does it lean on the previous section with words like “this”, “the above” or “said”?
- Check server logs. Do requests from OAI-SearchBot, GPTBot, PerplexityBot and ClaudeBot return successfully?
We covered the real role of structured data in this process in FAQPage, Schema and structured data: markup is a supporting layer and doesn’t close a gap in page text. We gathered all the structural reasons, from access to authority, in 7 structural reasons.
How does Recro make this gap visible?
The Recro Insight Model assesses a site’s technical state through the buyer’s question, not an abstract audit checklist: which information fails to reach the answer in which question, and which page gap causes it?
One simulation · 4 key pillars · 90+ steps
What sets Recro apart is four pillars built from scratch for each brand. Together they form a single simulation of 90+ steps, with each pillar producing the input for the next.
- Persona builder: Decision-maker profiles that represent the brand’s real buyers, specific to its sector and sales structure.
- Question builder: The brand-specific decision questions these people ask AI along their buying journey.
- Report builder: A report that reads the answers against the brand’s goals and competitors, together with the sources behind them.
- Action recommendation builder: A prioritized list of actions that closes the sector-specific signal gaps.
For page reading, each pillar produces the following:
- The persona builder establishes which decision-maker a page is critical for. A data sheet is critical to an engineer, a pricing model to a buyer.
- The question builder sets up the questions a page must answer. It shows whether headings match those questions.
- The report builder reads which of the brand’s pages are cited in answers and which information doesn’t carry into them.
- The action recommendation builder prioritizes page fixes within the sector-specific LLM signal list: access and first HTML first, then critical information with no text equivalent, then heading and section structure.
Sample finding format · illustrative
- Persona: Project engineer · Building materials · English
- Question: Elimination · Products with a given fire rating and test certificate
- Status: Product pages aren’t cited; a competitor’s technical page is quoted.
- Cause: Fire rating and test certificate exist only in a PDF sheet and a tabbed area loaded by JavaScript.
- Class: Access issue + Content GAP
- Action: Serve technical values as text in the first HTML; rewrite headings with class and standard names.
Executive summary
- AI systems read a page from raw HTML, split it into passages and match them to sub-questions; the visual experience is largely lost along the way.
- Seven gaps in human-first pages: JavaScript content, text on images, information locked in PDFs, vague headings, context-dependent paragraphs, missing numbers and conditions, access blocks.
- The design stays; the same information is also served in the first HTML, in plain text and under headings that answer the question.
Frequently asked questions
Do AI crawlers execute JavaScript?
According to Vercel and MERJ’s analysis, GPTBot, ClaudeBot and PerplexityBot don’t. Google’s systems can render pages with Googlebot’s infrastructure. Having key content in the first HTML is the safest approach.
Do we need to redesign our site?
Usually not. Most gaps can be closed with targeted fixes such as server-side rendering, moving information from images into text and rewriting headings.
Should we remove our PDF catalogs?
Keep them; they’re valuable as downloadable resources. But decision-critical information, capacity, certification, technical values, should also be in the HTML text of the page.
Does blocking GPTBot affect search visibility?
According to OpenAI’s documentation, GPTBot is used for model training and OAI-SearchBot for ChatGPT’s search index. You can refuse training and allow search; to be cited in ChatGPT search, OAI-SearchBot must not be blocked.
Does an llms.txt file close these gaps?
It doesn’t. Google has said it doesn’t use the file, and its effect in other tools isn’t proven. The real impact comes from the page itself being readable.
Does your site answer your buyers’ questions through a machine’s eyes? Recro builds a sample simulation with brand-specific personas and decision questions and shares the result in a demo report.
Request a demo insight report →
Sources
- OpenAI · Overview of OpenAI crawlers
- Vercel · The rise of the AI crawler (Vercel and MERJ)
- Search Engine Journal · Query fan-out technique in AI Mode
- Ahrefs · Why ChatGPT Cites One Page Over Another (Study of 1.4M Prompts)
- Cloudflare · Cloudflare just changed how AI crawlers scrape the Internet-at-large (2025)
- Google Search Central · AI features and your website
- Search Engine Journal · Google’s llms.txt guidance
This article was prepared with AI assistance and published under the review of Mehmet Semih İpek.



