DISPATCH · Nº 0552 · POINTCAST FIELD GUIDE
Who Gets to Read the Web?
A field guide to Firecrawl, spiders, indexes, and the argument hiding inside every fetch.
The web was born linkable. The spider made it legible. The index made it governable. AI made the old bargain visible.
Who Gets to Read the Web? is a PointCast field guide to Firecrawl and the larger system it enters: crawling, scraping, rendering, extracting, indexing, ranking, archiving, and deciding who may turn a public page into a portable record.
Firecrawl is an open-source web data API. Its main repository is AGPL-3.0, with separately identified MIT-licensed SDK and UI portions. It can search, map, scrape, crawl, render, interact with, and extract modern pages into Markdown, screenshots, links, or structured JSON. PointCast has no sponsorship, partnership, or financial relationship with Firecrawl.
A seven-stage browser-local instrument traces one URL through discovery, rules, fetch, render, extraction, indexing, and retrieval. It performs no network request, arbitrary execution, storage, analytics, account creation, or upload. An eight-position lens board turns the crawler through the perspectives of a publisher, librarian, researcher, builder, host, artist, agent, and public reader.
The history begins with the web's 1989-90 open linking architecture; the 1992 resource discovery problem; the 1993 World Wide Web Wanderer; the 1994 robots convention; the 1998 Google paper and PageRank; Beautiful Soup in 2004; Common Crawl in 2007-08; RFC 9309 in 2022; Firecrawl in 2024; and new publisher control and payment experiments in 2025-26. Primary and official sources are linked inline and carried in the JSON twin.
The feature's Crawl Compact asks a machine to identify itself, read the host's signals, minimize collection, keep a human pace, avoid access-control bypass, retain provenance, separate technical reachability from rights, and make stopping real.
This is an editorial and technical field guide, not legal advice. Scraping questions vary by facts and jurisdiction and can implicate copyright, privacy, contract, computer-access, database, publicity, consumer-protection, and sector-specific law. robots.txt communicates requested crawler behavior but is not access authorization or a security control.