Field scan · 2026-09-28
News and publishing
A first-hand scan of 16 sites and 10 repositories in news and publishing, measured on 2026-09-28. 15 of 16 sites name an AI crawler in robots.txt, 13 challenged or refused a scripted request, and 1 answered with a payment demand.
16 sites · 10 repositories · 10 findings · Markdown
Figures
Counts are out of every site scanned. A site that could not be read counts as not having the feature, so these are floors. The limits of the scan say which sites could not be read.
Findings
Measured means read first-hand, with the evidence kept. Inferred means a conclusion drawn from what was measured, and never above medium confidence.
- Measured · highNews publishers write the most detailed AI policies of any sector scanned. 15 of 16 robots.txt files name AI crawlers, and the average site blocks 14 of the 27 tracked crawler tokens outright. CNN blocks 26, The Verge 24, The New York Times 23. In developer tooling the figure was one site in eleven.
- Measured · highThe policy is enforced at the edge, not left to robots.txt. 13 of 16 sites challenged, refused or cut off at least one of four scripted requests. Only BBC, Wired, Substack served all four user agents the page, and two of those three disallow AI crawlers in robots.txt, so their block relies on the crawler's good behaviour.
- Measured · high8 of 16 sites gave at least one AI user agent a different answer than the Chrome user agent: The New York Times, The Washington Post, The Guardian, Financial Times, CNN, The Atlantic, The Verge, NPR. CNN answers all three AI user agents with HTTP 451, Unavailable For Legal Reasons. NPR closes the connection without an HTTP response.
- Measured · highTwo sites serve GPTBot and refuse ClaudeBot and PerplexityBot: The Guardian and The Verge. The Verge's robots.txt also allows GPTBot while blocking 24 other tracked tokens. Which crawler gets in is decided per operator, not per category.
- Inferred · mediumThe per-operator pattern matches publicly announced content licensing agreements between OpenAI and these publishers' owners. The scan cannot see a contract, so the link is an inference from who is served and who is refused.
- Measured · highOne HTTP 402 in the sector, and it cannot be paid. The Atlantic answers ClaudeBot with 402 and the message "Please contact the site owner for access": no price, no payment header, no x402 challenge. GPTBot and PerplexityBot get a 403 from the same site. No site in three sectors and 39 sites has offered an agent a price it could pay.
- Measured · highThe Atlantic is the only site declaring Content-Signal. It sets search=yes, ai-input=no, ai-train=no for four named crawlers: Googlebot, BingBot, DuckAssistBot and FacebookBot. They may index the page and may not use it in AI answers or training.
- Measured · highNo site publishes /llms.txt: 0 of the 8 where it could be requested. Publishers are telling agents what they may not do and are not describing themselves to them.
- Measured · highPublishing software is being written with agents while publishers block them. Across nine repositories and 5,526 commits in 90 days, 182 (3.3%) carry an AI-agent co-author trailer. Hugo leads at 39 of 177 (22%), every one signed by Claude. Gutenberg, the WordPress editor, has 118 of 2,074 (5.7%). The Guardian's own two repositories carry 13 agent co-authored commits, mostly Copilot, while its site refuses ClaudeBot.
- Measured · highWordPress core could not be measured. Its repository mirrors Subversion, credits contributors with a Props line, and carries no co-author trailers of any kind, so the method has nothing to read.
Sites scanned
The New York Times, The Washington Post, The Wall Street Journal, The Guardian, BBC, Reuters, Associated Press, Bloomberg, Financial Times, CNN, The Atlantic, Wired, The Verge, NPR, Medium, Substack.
Repositories scanned
Gutenberg, WordPress core, Ghost, Guardian dotcom-rendering, Guardian frontend, Newspack, Hugo, Jekyll, Eleventy (Build Awesome), Astro.
Method
Web: for each of 16 sites, fetched /robots.txt and /llms.txt once with a Chrome user agent, and the home page once each with Chrome, GPTBot, ClaudeBot and PerplexityBot user agents, 1.5 seconds apart, from one US address using a scripted HTTP client (no JavaScript, no cookies). Requests that returned no HTTP response were repeated once with a second client (curl); the 402 response was requested a second time to record its full headers. Each response was classed as served (200), challenged (a challenge page), refused (a 4xx with no challenge), payment required (402), or reset (no HTTP response). robots.txt was parsed into user-agent groups and checked against 27 tracked AI crawler tokens; a token counts as blocked when its group contains Disallow: /. Repositories: for each of 10 repositories, read every commit on the default branch from 2026-06-30 to 2026-09-28 through the GitHub commits API. A commit counts as AI co-authored when a Co-authored-by trailer carries an agent's name or an address an agent signs with, or when the commit author is a known agent account. A person's employer domain does not count. Automation accounts are counted separately.
Limits of this scan
Read these before quoting a figure.
- The web client sent a Chrome user agent string but is not a browser: it runs no JavaScript and has a different TLS fingerprint. A site that challenged it may serve a real browser normally. 'Challenged' describes what a script met, not what a reader meets.
- AI user agents were sent from an ordinary address, not from the operators' published ranges, and unsigned. A site that verifies crawlers by address or signature may admit the real crawler where it refused this one, or refuse it where it admitted this one. A site serving GPTBot's user agent from an unknown address is, if anything, trusting the string.
- Only the home page was requested. Article pages, which carry the content these policies protect, may be treated differently.
- robots.txt states a request, not an enforcement. A blocked token means the publisher asked that crawler to stay out.
- A token counts as blocked only when its own group contains Disallow: /. Partial disallows, and rules that reach a crawler through the * group, are not counted, so block counts are a floor.
- The 27 tracked tokens are the scanner's list, not an exhaustive one. Sites naming crawlers outside it are undercounted.
- Challenges and refusals left /llms.txt unreadable on eight sites. Those are unmeasured, not absent.
- WordPress core mirrors Subversion and carries no co-author trailers, so agent authorship there is unmeasured. Sector commit totals exclude it.
- Newspack has ten commits in the window and was archived in August 2026; its share is recorded at medium confidence.
- Co-author trailers are voluntary and undercount agent-written code where contributors strip them.
- One reading per site on one day, with a second reading only for requests that failed. Edge policies vary by region and change often.
How to cite
Overwing Atlas. "News and publishing: field scan of 16 sites and 10 repositories." Measured 2026-09-28. https://overwing.ai/atlas/scans/news-publishing-2026-09
Quote a figure with its date and its denominator. These are single readings on one day, and policies change. Headline findings are free to quote with attribution. Per-target rows are in the sector-scans dataset, under the Atlas data license.