Sources
What's in the archive
Whole public archives queryable with read-only SQL. Reddit, Hacker News, and Manifold Markets are documented in full; the wider corpus — scholarly literature, community archives, newsletters, social streams, prediction markets, reference works, public records, and code — is described by family.
Documented in full
reddit.comments · 29.5B+ posts + comments · Live The public Reddit archive — posts and comments queryable with read-only SQL by subreddit, author, timestamp, text, and score.Hacker News hackernews.items · 45M+ stories + comments · Live Every Hacker News story and comment, queryable with read-only SQL — source-native records, timestamps, authors, scores, and a live tail.Manifold manifold.markets The full Manifold prediction-market corpus — every market, every bet, every comment — queryable with read-only SQL and kept minutes-fresh from the live API.How coverage is stated
Each source page states what is covered: the public source, the SQL tables it lands in, the record shape, how fresh it is, and what is known to be missing — so an agent can tell the difference between "no results" and "not ingested". Known incompleteness is also served per relation as extent and known_holes in /v1/scry/schema.
The rest of the corpus
The full inventory
The corpus grows by more than a billion rows a day, and each relation states its own extent and known holes in the schema. The complete per-source inventory — counts, freshness, embedding coverage — is shared with evaluating customers.