# News and publishing: how the sector treats AI agents

> A first-hand scan of 16 sites and 10 repositories in news and publishing, measured on 2026-09-28. 15 of 16 sites name an AI crawler in robots.txt, 13 challenged or refused a scripted request, and 1 answered with a payment demand.

- Sector: News and publishing
- Measured: 2026-09-28
- Scan id: news-publishing-2026-09
- Publisher: Overwing Atlas (https://overwing.ai/atlas)
- Canonical page: https://overwing.ai/atlas/scans/news-publishing-2026-09

## Figures

| Figure | Value |
|---|---|
| Sites naming an AI crawler in robots.txt | 15 of 16 |
| Sites answering with a payment demand (HTTP 402, x402, pay per crawl) | 1 of 16 |
| Highest share of AI co-authored commits in one repository (gohugoio/hugo) | 22% |
| Repositories dormant since 2024 or earlier | 0 of 10 |
| Sites that challenged, refused or cut off a scripted request | 13 of 16 |
| Sites publishing /llms.txt | 0 of 16 |
| Sites declaring Content-Signal | 1 of 16 |

Counts are out of every site scanned. A site that could not be read counts as not having the feature, so these are floors. The limits of the scan say which sites could not be read.

## Findings

Each finding is marked measured (read first-hand, with evidence) or inferred (a conclusion drawn from what was measured).

1. **Measured, high confidence.** News publishers write the most detailed AI policies of any sector scanned. 15 of 16 robots.txt files name AI crawlers, and the average site blocks 14 of the 27 tracked crawler tokens outright. CNN blocks 26, The Verge 24, The New York Times 23. In developer tooling the figure was one site in eleven.
2. **Measured, high confidence.** The policy is enforced at the edge, not left to robots.txt. 13 of 16 sites challenged, refused or cut off at least one of four scripted requests. Only BBC, Wired, Substack served all four user agents the page, and two of those three disallow AI crawlers in robots.txt, so their block relies on the crawler's good behaviour.
3. **Measured, high confidence.** 8 of 16 sites gave at least one AI user agent a different answer than the Chrome user agent: The New York Times, The Washington Post, The Guardian, Financial Times, CNN, The Atlantic, The Verge, NPR. CNN answers all three AI user agents with HTTP 451, Unavailable For Legal Reasons. NPR closes the connection without an HTTP response.
4. **Measured, high confidence.** Two sites serve GPTBot and refuse ClaudeBot and PerplexityBot: The Guardian and The Verge. The Verge's robots.txt also allows GPTBot while blocking 24 other tracked tokens. Which crawler gets in is decided per operator, not per category.
5. **Inferred, medium confidence.** The per-operator pattern matches publicly announced content licensing agreements between OpenAI and these publishers' owners. The scan cannot see a contract, so the link is an inference from who is served and who is refused.
6. **Measured, high confidence.** One HTTP 402 in the sector, and it cannot be paid. The Atlantic answers ClaudeBot with 402 and the message "Please contact the site owner for access": no price, no payment header, no x402 challenge. GPTBot and PerplexityBot get a 403 from the same site. No site in three sectors and 39 sites has offered an agent a price it could pay.
7. **Measured, high confidence.** The Atlantic is the only site declaring Content-Signal. It sets search=yes, ai-input=no, ai-train=no for four named crawlers: Googlebot, BingBot, DuckAssistBot and FacebookBot. They may index the page and may not use it in AI answers or training.
8. **Measured, high confidence.** No site publishes /llms.txt: 0 of the 8 where it could be requested. Publishers are telling agents what they may not do and are not describing themselves to them.
9. **Measured, high confidence.** Publishing software is being written with agents while publishers block them. Across nine repositories and 5,526 commits in 90 days, 182 (3.3%) carry an AI-agent co-author trailer. Hugo leads at 39 of 177 (22%), every one signed by Claude. Gutenberg, the WordPress editor, has 118 of 2,074 (5.7%). The Guardian's own two repositories carry 13 agent co-authored commits, mostly Copilot, while its site refuses ClaudeBot.
10. **Measured, high confidence.** WordPress core could not be measured. Its repository mirrors Subversion, credits contributors with a Props line, and carries no co-author trailers of any kind, so the method has nothing to read.

## What was scanned

Sites (16): The New York Times, The Washington Post, The Wall Street Journal, The Guardian, BBC, Reuters, Associated Press, Bloomberg, Financial Times, CNN, The Atlantic, Wired, The Verge, NPR, Medium, Substack.

Repositories (10): Gutenberg, WordPress core, Ghost, Guardian dotcom-rendering, Guardian frontend, Newspack, Hugo, Jekyll, Eleventy (Build Awesome), Astro.

## Method

Web: for each of 16 sites, fetched /robots.txt and /llms.txt once with a Chrome user agent, and the home page once each with Chrome, GPTBot, ClaudeBot and PerplexityBot user agents, 1.5 seconds apart, from one US address using a scripted HTTP client (no JavaScript, no cookies). Requests that returned no HTTP response were repeated once with a second client (curl); the 402 response was requested a second time to record its full headers. Each response was classed as served (200), challenged (a challenge page), refused (a 4xx with no challenge), payment required (402), or reset (no HTTP response). robots.txt was parsed into user-agent groups and checked against 27 tracked AI crawler tokens; a token counts as blocked when its group contains Disallow: /. Repositories: for each of 10 repositories, read every commit on the default branch from 2026-06-30 to 2026-09-28 through the GitHub commits API. A commit counts as AI co-authored when a Co-authored-by trailer carries an agent's name or an address an agent signs with, or when the commit author is a known agent account. A person's employer domain does not count. Automation accounts are counted separately.

## Limits of this scan

Read these before quoting a figure.

- The web client sent a Chrome user agent string but is not a browser: it runs no JavaScript and has a different TLS fingerprint. A site that challenged it may serve a real browser normally. 'Challenged' describes what a script met, not what a reader meets.
- AI user agents were sent from an ordinary address, not from the operators' published ranges, and unsigned. A site that verifies crawlers by address or signature may admit the real crawler where it refused this one, or refuse it where it admitted this one. A site serving GPTBot's user agent from an unknown address is, if anything, trusting the string.
- Only the home page was requested. Article pages, which carry the content these policies protect, may be treated differently.
- robots.txt states a request, not an enforcement. A blocked token means the publisher asked that crawler to stay out.
- A token counts as blocked only when its own group contains Disallow: /. Partial disallows, and rules that reach a crawler through the * group, are not counted, so block counts are a floor.
- The 27 tracked tokens are the scanner's list, not an exhaustive one. Sites naming crawlers outside it are undercounted.
- Challenges and refusals left /llms.txt unreadable on eight sites. Those are unmeasured, not absent.
- WordPress core mirrors Subversion and carries no co-author trailers, so agent authorship there is unmeasured. Sector commit totals exclude it.
- Newspack has ten commits in the window and was archived in August 2026; its share is recorded at medium confidence.
- Co-author trailers are voluntary and undercount agent-written code where contributors strip them.
- One reading per site on one day, with a second reading only for requests that failed. Edge policies vary by region and change often.

## How to cite

Overwing Atlas. "News and publishing: field scan of 16 sites and 10 repositories." Measured 2026-09-28. https://overwing.ai/atlas/scans/news-publishing-2026-09

Quote a figure with its date and its denominator. These are single readings on one day, and policies change.

## Data

- Headline findings, free: https://overwing.ai/api/v1/atlas/summary
- Per-target rows (Atlas Pro or Team): https://overwing.ai/api/v1/atlas/datasets/sector-scans
- License: https://overwing.ai/atlas/license
