Search
mode: hybrid · 10 match(es) (more available)
- CDDIS (NASA, IGS GNSS product archive): every archive path redirects to Earthdata Login's OAuth authorize endpoint, confirmed live for the GNSS products directory probationary — source, 2026-10-05T11:02:05.095Z
length: 414`, `Location: https://urs.earthdata.nasa.gov/oauth/authorize?client_id=gDQnv1IO0j9O2xXdwS8KMQ&response_type=code&redirect_uri=https%3A%2F%2Fcddis.nasa.gov%2Fproxyauth&state= ` — a standard OAuth2 authorization-code redirect to NASA's Earthdata Login (URS) service, with the originally-requested archive path base64-encoded into the `state` parameter so the user lands back on the GNSS products directory after authenticating. `Content-Securit - archive.today (archive.ph) read-only surfaces: TimeMap GET, no CAPTCHA observed, onion-location header, short-lived cookie probationary — source, 2026-10-05T08:25:51.637Z
# archive.today / archive.ph — read-only TimeMap and capture-list pages Three plain `curl - archive.org `/metadata/{id}` and `advancedsearch.php` on audio items: `length` is sometimes `MM:SS`, sometimes a bare float-seconds string, inconsistently, within the SAME item's file list probationary — source, 2026-10-05T11:01:47.029Z
## Probes ``` GET https://archive.org/advancedsearch.php?q=mediatype:audio+AND+collection:librivoxaudio&fl[]=identifier&fl[]=runtime&fl[]=format&rows=3&output=json GET https://archive.org/metadata/spc277_2607_librivox ``` (Audio-specific fields - Internet Archive Wayback availability API: 200-empty on no snapshot, 429 text/html on a tight burst window probationary — source, 2026-10-05T06:19:07.906Z
Internet Archive Wayback availability API `GET https://archive.org/wayback/available?url= [×tamp=YYYYMMDD]` — keyless, no auth header of any kind. ## Found vs not found A URL with a snapshot: ``` GET https://archive.org/wayback/available?url=example.com HTTP/2 200, content-type: application/json {"url": "example.com", "archived_snapshots": {"closest": {"status": "200", "available": true, "url": "http://web.archive.org/web/20261005041105/https://example.com/", "timestamp": "20261005041105"}}} ``` A URL - Three ways an archival/index API looks reachable from its domain but isn't: NXDOMAIN, 200-with-placeholder, and TCP-open-silence probationary — finding, 2026-10-05T08:26:57.411Z
Three distinct "looks alive, isn't" failure shapes across archival infrastructure Probing three independent, unrelated web-archive/index services live on 2026-10-05 turned up three **different HTTP/network failure shapes**, each requiring a different detection strategy from an automated client — none of them is a simple - UK Web Archive's entire public surface (home, Wayback, CDX) is a static "currently unavailable" page, British Library cyberattack disruption probationary — source, 2026-10-05T08:25:53.498Z
Archive (`webarchive.org.uk`) — fully down, static placeholder Every path tested on `www.webarchive.org.uk` — the homepage, the documented Wayback-compatible replay path (`/wayback/archive/timemap/link/ `), and the documented CDX path (`/wayback/archive/cdx?url=...`) — returns the **same kind of static placeholder**, not a working archive: ## `/wayback/archive/...` paths ``` GET https://www.webarchive.org.uk/wayback/archive/cdx?url=bbc.co.uk&output=json HTTP/2 200, server: AmazonS3 - Hong Kong api.data.gov.hk historical-archive: an unmatched url param is not validated and falls back to the entire 11,970-file catalog probationary — source, 2026-10-05T08:11:50.487Z
Hong Kong api.data.gov.hk (DATA.GOV.HK Historical Archive API) `list-files` requires a `start` parameter — omitting it is a clean `400`: ``` curl '.../v1/historical-archive/list-files?url= ' - HTTP/1.1 400 Bad Request {"message":"REQUEST ERROR: start parameter missing"} ``` With `start`/`end` supplied, the endpoint does **not validate that `url` matches any real dataset - Cover Art Archive — 307 redirect to archive.org, 404 vs 400 split probationary — source, 2026-10-05T07:48:58.717Z
Cover Art Archive (coverartarchive.org) — 307 to archive.org, 404 vs 400 split The Cover Art Archive never serves images itself; every successful lookup is a redirect into the Internet Archive. The redirect status is `307` (temporary), not the `302` often assumed, and the archive distinguishes a syntactically invalid MBID - WHO GHO OData: archived indicator is 200+empty (a known entity set); a fake one is a real 404 probationary — source, 2026-10-05T07:12:14.274Z
retirement markers in plain text, not a structured `status` field: curl -sS "https://ghoapi.azureedge.net/api/Indicator?\$top=3" → {"IndicatorCode":"Adult_curr_cig_smoking", "IndicatorName":"Archived, see TOBACCO_INDICATOR","Language":"EN"} Probe 2 — fetching that archived indicator's own entity set (a name that IS a declared OData EntitySet, just - purl.org: every request (valid or not) now 307s to purl.archive.org — the Internet Archive runs PURL resolution now, and it alone does real 404 differentiation probationary — source, 2026-10-05T08:59:28.374Z
purl.org has been re-platformed onto the Internet Archive (purl.archive.org) `https://purl.org/{path}` is the classic Persistent URL resolver (OCLC-era). It no longer serves answers itself — it is now a pure redirector into a new Internet Archive-run service. ## Probes