Asosiy kontentga o'tish
BiznesMarketing
Sign in
MarketingAugust 18, 2026· 2 min read

Which Pages Does ChatGPT Actually Read? RESONEO Analyzed 1,200 Answers

Which Pages Does ChatGPT Actually Read? RESONEO Analyzed 1,200 Answers

In July, the French agency RESONEO analyzed 1,200 ChatGPT responses, 88,000 search results, and 26,900 unique pages to figure out which part of the internet the chatbot actually draws on when generating answers. The study found that ChatGPT's answer-generation system has three layers: an index that discovers pages (a search catalog), a cache that stores full copies of pages (temporary memory), and a small set of pages the model opens and reads live.

Until July, researchers had been tracking a field called result_source in the data stream ChatGPT sends to the browser — it listed four internal pipelines named labrador, bright, oxylabs, and serp, though OpenAI has never officially acknowledged their existence. Starting July 21, that field abruptly vanished across every account being monitored.

The most striking finding is that free and paid ChatGPT users draw on entirely different slices of the internet. In the free "Think" mode, 74.7% of results come from OpenAI's own internal index, labrador, with only 3.1% pulled from classic Google search results. In the paid "Thinking" mode, it's the reverse: 75.3% of results come from scraping Google, and only 24.7% from labrador. In other words, two users asking the exact same question can end up with answers built from completely different sources.

The study also exposed ChatGPT's caching mechanism: once a page is read, the model converts it from full HTML into Markdown and stores it — and that stored copy is shared across all users. If a free user in Berlin asks a question, they get the same cached copy of a page that a paid user in Ohio had requested a week earlier. The copy is considered "fresh" for roughly 30 minutes, after which the old version keeps being served while a refresh happens in the background — and a site's Cache-Control: no-store or noindex directives are simply ignored.

The study's authors note that in "instant" mode, the only "foundation" the user actually sees is roughly the first 200 characters of text beneath the main headline (H1) — the meta description is ignored entirely. That makes writing clear, substantive opening sentences right after the H1 practically important for sites, even as these systems keep changing rapidly.

Source: Search Engine Land · view original article
Ulashish:TelegramLinkedIn
← Back to homepage

Related articles