Aggregate data · version 1 · schema 1.3.0
BooksReddit book-mention dataset
Published aggregate book-mention statistics from the Reddit communities tracked by BooksReddit, with reconciled breakdowns, source-window coverage, and sentiment context from displayed excerpts.
What is included
- 505 published book records.
- Canonical mention totals and counts of release-wide, case-insensitively distinct Reddit account identifiers.
- Sentiment context from the source excerpts displayed on book pages.
- Complete monthly, per-subreddit, and subreddit-by-month breakdowns. Every exported margin is checked against the canonical total before the build can finish.
- A privacy-safe source coverage ledger for each bounded partial unit: source filename, subreddit, requested and covered windows, score floor, page cap, selected count, and recovery evidence.
- No Reddit usernames or account identifiers, raw comments, commenter email addresses, or other commenter-level personal data. Published book records do include author names.
Source coverage in this release
The release selected 324
subreddit-window units from 305
source result files. 324 cover their complete requested windows; 0 are bounded newest-first partials. The JSON download exposes those partials under
source_coverage.partial_units without publishing comments, excerpts, usernames,
or local cache paths. Each row also reports whether the current selected snapshot is pinned
in the reviewed manifest; a pin makes that current selection reproducible but does not prove
it matches a historical partial.
0 selected legacy result files predate explicit result-level coverage metadata, and 280 selected uncapped legacy caches do not have completeness sidecars. Those complete-cache windows rely on the documented namespace invariant; every known capped sample is kept separate and listed below. complete legacy cache artifacts rely on the invariant that partial walks were never cached; capped samples require a separate namespace and verified incomplete-coverage sidecar
- 0 count-matched legacy recoveries. Matching the legacy comment count confirms count equality only; it does not establish historical content equivalence with the original partial archive.
- 0 reviewed mismatch snapshots. A reviewed accepted mismatch pins repeatedly confirmed, stable current archive content; equivalence with the historical legacy partial remains unknown.
How to cite it
Cite “BooksReddit aggregate book mentions, version 1,” link to this page, and include the
download's generated_at and data_updated_at values. The first is
the export build time; the second is the latest included record update. Counts may change
as new source batches are processed or matches are corrected.
Method and limitations
These are descriptive aggregates of the communities tracked by BooksReddit, not a poll of all Reddit users or all readers. Read the full methodology for extraction, ranking, sentiment, and coverage details.
- The dataset describes a score-filtered sample of tracked communities, not all Reddit comments or all readers.
- A manifest-selected source unit can cover its full requested window or a bounded newest-first partial sample. Coverage varies by unit and release; the export's global first and last month do not imply uniform coverage.
- Matching the legacy comment count confirms count equality only; it does not establish historical content equivalence with the original partial archive.
- A reviewed accepted mismatch pins repeatedly confirmed, stable current archive content; equivalence with the historical legacy partial remains unknown.
- 0 selected legacy result files predate explicit result-level coverage metadata, and 280 selected uncapped legacy caches lack completeness sidecars; those cache windows rely on the documented complete-cache namespace invariant.
- All published count breakdowns reconcile exactly within each record, but only for the source artifacts and covered ranges included in that release.
- The unique_commenters field deduplicates case-insensitive account identifiers across the selected release corpus; it is not a count of people.
- Average sentiment summarizes scored displayed source excerpts, not every mention, and is null when none are scored.
Provenance, schema, and reuse
The aggregates derive from public Reddit discussion accessed through Arctic Shift . The JSON Schema defines every exported field and CSV column. No dataset-specific reuse license is currently declared; do not assume the site's software license applies to source-derived data. Account identifiers are used only to compute distinct counts and are not included in the download.