nextbookinorder.com · 2026 to now
A bug that ships once gets a deploy check
A reading-order catalogue for book series, more than 68,000 pages, that I built and run end to end. Its LLM pipeline writes nothing when a page has no usable source, and my rule is that each class of bug that ships gets a check in the deploy gate.
- Role
- Founder and Engineer of an independent product, 2026 to now
- Scale
- More than 68,000 pages. About 150M rows staged in PostgreSQL.
- Built with
- Python, PostgreSQL, local open-weight LLMs, Astro, S3, CloudFront, Lambda for search
- Live at
- nextbookinorder.com
What it is
nextbookinorder.com answers one question: what order do I read this series in? It is a static site of more than 68,000 pages, generated from a PostgreSQL catalogue.
The pipeline
The catalogue starts as two bulk sources: a 144 GB compressed dump and a second source of 62M records. A streaming extraction filters them, and about 150M rows are staged in PostgreSQL. Entity resolution then links records across the two sources. A much smaller catalogue comes out the other side.
Matching runs in confidence tiers. A shared identifier is a strong match. A name plus a birth year is a weak one. It stays pending until it is adjudicated, and unsure pairs stay unlinked. When two candidates fit, neither is picked.
A static site gets no second chance
No server renders these pages. A wrong page stays wrong until the next build and deploy, so the checks have to run before the upload.
An LLM pipeline that is allowed to write nothing
Some of the prose on the site is generated by a local open-weight model. The standard for new prose has six stages: a source pack, generation, deterministic filters, a judge, a claim audit in code, and storage with provenance. The author-biography and series-premise generators run all six. The book-synopsis writer has a grounding check and a judge that sees only its source, and some older text predates the standard. The judge is a different model from the generator. Today both are Qwen3 checkpoints, Instruct to write and Thinking to judge. The rule at every stage is grounded or skip, so a page with no usable source stays short.
I had a separate agent workflow measure where this fails. It sampled about 60 pages from one pilot, and about 18% of the generated summaries on thin sources contained fabrications. A model judge alone had not been enough. The fix was a stricter gate in the generator, which now refuses a source that does not describe the book itself and writes nothing for it, together with that pilot's judge from a different model family and a deterministic faithfulness check. A fresh sample then measured 3.3%. Both rates are small-sample estimates.
One failure shipped. Of the 393 series pages carrying a generated About section, 227 had a paragraph that restated the intro above it. The generator had been given the intro in its source pack, and the judge never saw the page the paragraph would join. I removed the section and, where a judge passed the result, merged its unique sentences into the intro. The generators that came afterwards for author biographies and series premises show their judge the rest of the page.
The deploy gate
My rule is that every class of bug that ships once gets a check in the deploy gate. An error from any check aborts the deploy. The gate runs after the build and before anything uploads. The deploy script has an emergency bypass, an environment variable. My rule is never to use it.
- The gate grew from 29 check functions on August 1, 2026 to 63 on September 30, with 45 ratchet baselines.
- Baselines move only with a written reason. A ceiling comes down when a cleanup lands. It goes up only after the new rows behind the change are isolated and counted.
- Template faults are sampled with a fixed seed. Data faults get a full scan: one deploy checked 223,000 outbound product links and found none outside the verified manifest or the pre-order exemption.
- On the September 30 build, eight self-tests ran first. They cover the link, region and cover guards, mostly by feeding them broken input, so those checks can't pass everything behind a green log.
When a ratchet regresses, I look for a bug before I touch the baseline. In August one count rose by 83. Those 83 were series about to be exported with volumes missing.
The gate has had bugs of its own. One check sampled from an unordered set, so the same input passed at 23:28 and failed at 00:11. It scans everything now.
How the code gets written
I build this product with AI coding agents. I own the architecture, the rules and what the gate must catch. Agents write most of the code, the gate's included. The self-tests above keep a broken check from passing behind a green log. That practice has its own write-up: running AI coding agents under independent review.