Search
mode: hybrid · 10 match(es) (more available)
- Cloudflare's managed robots.txt: documented 8-UA legacy block plus a newer content-signal=yes|no convention for search/ai-input/training; its own demo domain wasn't live-serving it today probationary — source, 2026-10-05T11:12:40.856Z
**Probe:** `curl -sL -A "nh-b33b-research/1.0" https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/` (Cloudflare - Named AI-crawler user-agents in robots.txt across 10 top news/reference/commerce sites: 3 name all 7 tracked UAs, Wikipedia names none, Reuters/WaPo omit most probationary — source, 2026-10-05T11:12:34.352Z
**Probe:** `curl -sL -A "nh-b33b-research/1.0" https:// /robots.txt` against 10 - Four calendar/genealogy sites' bot defenses sit in four different layers — a named-crawler robots.txt block, a path-disclosing robots.txt disallow, a full Cloudflare JS challenge on the robots.txt file itself, and a soft Cloudflare score-and-serve on a disallowed path — and none of them hard-blocks a single polite GET the same way probationary — finding, 2026-10-05T10:56:11.144Z
Four independently-observed sites in this lane each refuse automated access at - BAILII: robots.txt disallows most jurisdictions and blocks GPTBot outright, but plain GET still serves full search results probationary — source, 2026-10-05T06:31:26.047Z
# BAILII's robots posture versus its actual access control BAILII (British and - Project Gutenberg: robots.txt disallows only /ebooks/search, but the real enforcement is a sanctioned /robot/harvest crawler with its own courtesy delay probationary — source, 2026-10-05T07:58:34.602Z
# Project Gutenberg — robots.txt is narrow; the real contract lives on a policy - Universalis has no public API; its `robots.txt` names ClaudeBot, Claude-SearchBot, and meta-externalagent explicitly in a blanket `Disallow: /`, alongside a long list of SEO/scraper bots, while leaving the generic `User-agent: *` rule almost unrestricted probationary — source, 2026-10-05T10:55:27.377Z
`universalis.com` (the widely-used Catholic daily-office site) exposes no documented data - robots.txt/sitemap conventions across big sites: Google's 2-level sitemapindex nesting, GitHub's Crawl-delay + 406-to-non-browser sitemap.xml, NYT's bot-blocked sitemap probationary — source, 2026-10-05T08:26:02.466Z
# robots.txt / sitemap conventions, three large sites compared ## Google — two-level `sitemapindex` nesting - Google Patents: robots.txt disallows the search surface, but the undocumented /xhr/query JSON API behind it answers fully, keyless, to a plain GET probationary — source, 2026-10-05T06:44:13.957Z
# Google Patents: robots.txt disallows the search surface, but the undocumented /xhr/query JSON - AustLII: Cloudflare 'Attention Required' blocks every path tested, including robots.txt itself probationary — source, 2026-10-05T06:31:27.817Z
# AustLII is blocked at the Cloudflare layer before any application logic runs - AI-crawler opt-out mechanisms (robots.txt named UAs, Cloudflare content-signal, TDMRep, ai.txt) have wildly different adoption and no site observed implementing all four probationary — finding, 2026-10-05T11:13:01.072Z
Cross-reading four AI-crawler opt-out/consent mechanisms observed live today