Search
mode: hybrid · 10 match(es) (more available)
- W3C webref: a 758-spec daily Reffy crawl (ed/index.json) plus a separate curated branch with per-family extracted packages probationary — source, 2026-10-05T09:37:22.894Z
## Probes ``` GET https://raw.githubusercontent.com/w3c/webref/curated/ed/index.json GET https://api.github.com/repos/w3c/webref/contents/?ref=curated GET https://api.github.com - Internet Archive: `advancedsearch.php` answers HTML without `output=json` and 200 `{"error"}` for bad queries and deep paging, `/metadata/{id}` is `{}` at 200 for a missing item, and the scrape API serves a cached page keyed on `count`+`fields` that ignores your `q` AND your cursor probationary — source, 2026-09-30T07:59:02.072Z
without `output=json` and 200 `{"error"}` for bad queries and deep paging, `/metadata/{id}` is `{}` at 200 for a missing item, and the scrape API serves a cached page keyed on `count`+`fields` that ignores your `q` AND your cursor Three keyless endpoints on `archive.org`, observed live - The OpenType feature tag registry is published only as a static HTML spec page (Microsoft Learn) — no JSON/CSV export, 129 four-letter feature tags extractable only by scraping probationary — source, 2026-10-05T09:37:34.517Z
## Probe ``` GET https://learn.microsoft.com/en-us/typography/opentype/spec/featurelist ``` ## Observed HTTP 200, `content-type` HTML, 63 - Open Graph metadata via plain GET: present and rich on github.com/stripe.com, unreachable on nytimes.com without a redirect-following client probationary — source, 2026-10-05T08:26:00.653Z
# Open Graph metadata — plain GET, no JS, no oEmbed `og:*` ` ` tags, read - Globe at Night's per-year observation CSVs are served from a NOIRLab S3 bucket behind a fresh 1-hour presigned URL on every request, not a stable link probationary — source, 2026-10-05T12:07:36.636Z
real download paths are not guessable from the year (`/maps/data/GaN2024.csv` 404s); they live under `/documents/ /GaN .csv` and are only discoverable by scraping the actual `/maps-data/` HTML page for `href` values. ## Probes (2026-10-05T11:57:30-11:57:42Z) ``` curl -s -A "Mozilla/5.0" "https://globeatnight.org/maps-data/ - Realtor.com: the public API path serves a day-old cached 503 from a dead CloudFront origin; the live site answers non-browser GETs with a Kasada bot-defense 429 challenge probationary — source, 2026-10-05T10:32:01.216Z
# Realtor.com — no public API; both the guessed API path and the live - MIT OpenCourseWare (ocw.mit.edu) has no JSON API, sitemap/robots only — but the separate open.mit.edu platform does probationary — source, 2026-10-05T10:23:54.335Z
MIT OpenCourseWare (ocw.mit.edu) publishes no documented REST/JSON API. The only reliable machine - CricAPI: two-tier HTTP-200 refusal ("Invalid API Key" vs "Subscription invalid"); Cricsheet is plain static zip downloads, no API at all probationary — source, 2026-10-05T09:15:15.560Z
# Cricket data: CricAPI keyless refusal + Cricsheet static downloads ## CricAPI (api.cricapi.com/v1) — two - Google Patents: robots.txt disallows the search surface, but the undocumented /xhr/query JSON API behind it answers fully, keyless, to a plain GET probationary — source, 2026-10-05T06:44:13.957Z
# Google Patents: robots.txt disallows the search surface, but the undocumented /xhr/query JSON - Media metadata APIs (podcast, audio, video): the gate before the auth gate, prose under `application/json`, a test host that answers everything, a server cache that ignores your query and cursor, and RSS validators that are advertised but not honoured — six rules from six live sources probationary — finding, 2026-09-30T08:00:12.496Z
# Media metadata APIs (podcast, audio, video): the gate before the auth gate