The problem
My coworker Bea wanted to know who reads what in Spain, especially young readers. She looked for a dataset that could answer it and couldn’t find one. That question started this project.
The data exists, but it’s spread across population estimates, reading surveys, library statistics, bestseller lists and large book catalogues. Put side by side, they look like they should add up to an answer. They don’t, because they don’t describe the same people or measure the same thing. So before estimating anything about readers and titles, I set out to make the sources comparable without hiding where they differ.
The sources
The FGEE reading survey reports young readers as one 14–24 band. INE publishes population by single year of age. My target group is 15–24. I can build that population from INE, but I can’t take 14-year-olds out of FGEE’s estimate without survey detail I don’t have. So the two bands stay as they are, and any comparison says so.
Bestseller lists give positions within a bookshop panel: no sales figures and no reader age. Library statistics count loans, also without age. And buying, borrowing and reading a book are three different things.
Then there’s the question of what “a book” is. An ISBN identifies an edition. A reader thinks about the work, across formats and translations. Matching editions to works is where joins quietly multiply rows or pick the wrong record, so the matching code refuses ambiguous joins instead of guessing.
Edition → work
Each ISBN identifies an edition candidate. Reviewed links can connect editions to one work; ambiguous matches stay unresolved.
Architecture and decisions
The pipeline follows raw → bronze → silver → gold:
- Raw keeps the source files exactly as downloaded.
- Bronze parses them into typed Parquet.
- Silver cleans each source and conforms the population and survey fields.
- Gold holds analytical tables with explicit definitions.
Python and Polars do the transformations; DuckDB handles queries and joins. Population counts, survey estimates, service metrics and rankings each keep their own grain instead of being forced into one table.
The sources don’t all have the same reuse terms, so the project has two profiles. The public profile uses statistics I can redistribute: INE, eBiblio, and a small set of FGEE figures with attribution. Goodreads (via the UCSD Book Graph), Amazon metadata and the bestseller captures stay in local research only.
Between the two sits a publication gate. It’s deterministic and denies by default. Before anything is published, it checks file hashes, the reviewed rights policy and every upstream dependency, including filters and matching decisions. If an output’s lineage touches a restricted source, it’s blocked. Dropping the restricted columns or averaging them away doesn’t get around it.
Two source tiers
Public: INE, eBiblio and selected FGEE figures. Restricted local research: Goodreads/UCSD, Amazon and bestseller lists.
The deterministic gate blocks any artifact whose lineage touches restricted sources. The public sample builds independently and produces a local approved bundle.
Every run writes a manifest with source versions, schemas, row counts, checksums, validation results and timings. The gate writes a receipt naming the exact artifacts it approved.
What it does today
The public sample runs offline in about a minute: 22 INE population observations, 10 eBiblio metrics and 39 selected FGEE reading and purchasing figures. It builds three gold tables, and the gate checks 14 artifacts before approving the bundle. The release branch has 41 passing tests and green CI covering both profiles, the gate, formatting and chart reproduction.
The full local run processes about 6.07 million rows from 8.2 GB of raw data: 4,448,181 Amazon metadata rows, 1,459,609 Goodreads rows across eight genre files, 165,640 INE rows, 398 bestseller-list rows, 10 eBiblio metrics and 3 survey indicators. These are rows at different grains, not unique books or readers.
The code will be public soon.
Limits and what’s next
The piece Bea actually asked about, which titles each age group reads, is still missing. Goodreads and Amazon don’t record a reader’s age or country, and neither a genre label nor an edition’s language can stand in for it.
Two things could close the gap: an FGEE breakdown of age by title, with its definitions and margins of error, or a survey of my own that asks both age and titles. In parallel I’m working on bounded ingestion, per-row lineage and reviewed edition-to-work matching.