Files
bggpipe/docs/spec.md
T
Eric WagonerandClaude Fable 5 c7fdc60f87 Audit round 8: the export earns its publishing promises
Five blind reviewers over the day-old export stage; ~30 findings
verified, the big ones sharing one root — a static-site generator
makes promises a pipeline stage doesn't, and the first cut kept none
of them.

URL stability was empirically false two ways: adding an edition whose
key sorted first STOLE the base slug (every colliding URL reshuffled),
and removing the base holder renumbered survivors over the stale
pages' corpses — wrong content at live URLs, not even 404s. And
nothing ever deleted anything: removed games stayed published forever.
One mechanism fixes all of it — a manifest (.bggpipe-export.json) in
the output directory records which slugs the export owns and which
source URL produced each cover. Slugs persist across runs (a published
URL never moves and can never be stolen), stale pages are removed
(only ever manifest-claimed ones — user files are not ours to touch),
replaced box art re-fetches when its URL changes, and "art" is a
reserved name so a game called Art can't move into the asset dir.

Trust-the-network fixes: a 200 response must LOOK like an image (magic
bytes + size) before it's cached, else a CDN interstitial became a
permanent "cover" that re-runs skipped forever; downloads go through
fsio.atomic_write_bytes instead of a hand-rolled fixed-tmp-name dance
(the exact hazard fsio's own docstring warns about); a missing
hand-added cover counts as a failure instead of silently shipping
coverless; the badge file is sniffed too; CDN pacing raised to 1s and
written into the spec as an adjudicated carve-out rather than a code
comment's private opinion.

Ship-shape: pages write atomically with the index LAST (a killed run
can't publish links to pages that don't exist); the CLI exits nonzero
on failures so `export && rsync` can't publish an incomplete site;
footer/fine-print contrast now clears WCAG AA on the sky background;
meta description, og:title/og:image and a favicon stop bare unfurls;
the BGG link moved out of the h1; the noart tile is aria-hidden; the
search box gained a no-matches message; numeric fields from enrich
render instead of crashing the join; years and ids are escaped; the
players/playtime formatters are aligned with their JS twins and both
sides carry keep-in-sync constraint comments; export moved after
enrich in the CLI listing.

Twelve export tests now, including the previously-vacuous atomicity
test rebuilt to actually interrupt a write. One honest loose end: one
full cover re-fetch occurred during rollout that the identical naming
code can't explain; the manifest's URL records make any recurrence
diagnosable. 352 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jXZFSTZQKzAC8fqpWSz9g
2026-08-06 18:51:05 -04:00

157 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Shelf-to-BGG Collection Pipeline — Design Contract
> This is the pipeline's contract — what must stay true, and why — kept
> current as the design evolves. Read it before changing pipeline
> semantics. How to *use* the pipeline is the [user's guide](guide.md).
## Goal
Build a command-line pipeline that takes photos of my board game shelves and ends with every recognized game added to my BoardGameGeek collection. The pipeline must be resumable, idempotent, and leave a human-reviewable audit trail at every stage.
Beyond the bare game, capture **which edition/version I own** wherever the photos allow it — many of my games exist in multiple editions, and in some cases I own more than one edition of the same game (each is a separate collection entry). Also capture **all critical data about each game itself** (see Stage 6 — enrich): the BGG `/thing` metadata is cheap to fetch and will seed the future frontend. Purchase provenance (where/when acquired, price paid) is explicitly NOT tracked — I don't have that data.
**Out of scope (for now):** the web frontend for displaying the collection. That's a follow-up project. But keep the data artifacts (see Data Model) clean and structured so they can seed it later.
## Constraints & Context
- BGG has **no write API**. Reads go through the XML API2 (`https://boardgamegeek.com/xmlapi2/`); writes must automate the website itself with a real login session.
- The XML API **requires a registered application** (boardgamegeek.com/using_the_xml_api, policy 2025-07-02): every request carries `Authorization: Bearer <token>` from the `BGG_API_TOKEN` env var, sent to `boardgamegeek.com` without a leading `www`. Register a free non-commercial application at boardgamegeek.com/applications — approval can take a week+; offline development and tests run on recorded XML fixtures.
- BGG's XML API queues collection requests: a first call may return HTTP 202 ("try again"). Retry with backoff.
- BGG will throttle aggressive clients. Target ≤1 request every 2 seconds to any BGG endpoint, with jittered backoff on 429/503.
- Image-CDN carve-out (adjudicated 2026-08-06): cover downloads from BGG's image CDN (`cf.geekdo-images.com`) during `export` pace at ≥1 second apart, fetch only files not already on disk, and re-fetch a file only when its source URL changes. The CDN is not the API, but it is still BGG's infrastructure — the spacing is mandatory, not advisory.
- BGG changed API access policies in 2025; some older community tools broke. Don't depend on undocumented endpoints beyond XML API2 and the public website.
- Vision extraction uses the **Anthropic API** (Claude with vision). Assume `ANTHROPIC_API_KEY` in the environment.
- Runs on macOS. Prefer **Python 3.12+** with `uv` for dependency management. Browser automation via **Playwright** (not Selenium).
## Pipeline Overview
Five stages, each a separate subcommand of one CLI (suggest `bggpipe`):
```
photos/ → [1 extract] → titles.json → [2 resolve] → matches.csv
→ [3 review] → matches.csv (approved) → [4 diff] → to_add.csv
→ [5 upload] → upload_log.csv
→ [6 enrich] → games.json (runs any time after review)
```
Each stage reads the previous stage's artifact and writes its own. Re-running a stage must be safe (skip already-processed items).
### Stage 1 — `extract`: Vision title extraction
- Input: a directory of shelf photos (JPEG/PNG/HEIC — convert HEIC to JPEG first via `sips` or Pillow-HEIF).
- For each photo, send to the Anthropic API (model: latest Sonnet) with a prompt that asks for:
- Every board game title visible (spines and face-out boxes), transcribed as printed.
- A confidence level per title (`high` / `medium` / `low`).
- **Edition/version cues**, each if legible: publisher name or logo, edition wording ("2nd Edition", "Deluxe", "Big Box", anniversary marks), copyright/print year, language, and distinctive box-art notes (colorway, artwork style). These drive version matching in Stage 2.
- If the SAME title appears more than once in the photos with visibly different boxes, report each as a separate entry — I own multiple editions of some games.
- Downscale images so the long edge is ≤1568px before sending (API sweet spot; keeps tokens down).
- Prompt for structured JSON output; parse defensively (strip code fences).
- Handle overlap: the same game may appear in two photos. Dedupe by normalized title (casefold, strip punctuation/articles) but **keep provenance** — record which photo(s) each title came from.
- Output: `titles.json` — list of `{title_raw, title_normalized, confidence, publisher_hint, edition_hint, year_hint, language_hint, art_notes, source_photos[]}`.
- Dedupe nuance: identical normalized titles from different photos collapse to one entry ONLY if their edition cues don't conflict; conflicting cues (different publisher/edition text) stay as separate entries.
- Support `--only <photo>` to re-run a single photo (e.g., after retaking a blurry shot).
- **Unidentified sightings**: boxes that appear to be games but can't be confidently titled (blurry, obscured, sharp angle, frame edge) are reported rather than silently omitted — location described relative to identified neighbors, plus any partial text and art notes → `unidentified.json`, keyed by photo. The end-of-run summary lists them (and low-confidence reads) so I can take a closer photo and re-run with `--only`.
### Stage 2 — `resolve`: Match titles to BGG IDs
- For each title, query `https://boardgamegeek.com/xmlapi2/search?query=<title>&type=boardgame,boardgameexpansion`.
- Scoring heuristic for candidates:
1. Exact normalized-name match → strong.
2. Fuzzy match (e.g., `rapidfuzz` token_sort_ratio ≥ 90) → good.
3. If multiple candidates tie, fetch `/thing?id=...&stats=1` for the top ~5 and prefer higher-owned/higher-ranked entries (obscure duplicates lose to the well-known game of the same name).
- Classify each result:
- `auto` — single confident match, no review needed.
- `ambiguous` — multiple plausible matches; store all candidates.
- `unmatched` — nothing plausible found.
- Expansions: BGG returns `boardgameexpansion` as a distinct type. Keep them — I own expansions and want them in the collection — but tag them so review can catch base-game/expansion confusion (a spine reading "Wingspan Europe" must not match base Wingspan).
- Cache all BGG responses on disk (keyed by query/ID) so re-runs don't re-hit the API.
- **Post-resolve dedupe**: rows resolving to the same (bgg_id, version_id — or both version-unknown) are the same physical game read twice (typo, partial spine) unless their extraction cues conflict (two editions). Losers get `match_status=merged` + a `merged_into` column pointing at the survivor — rows never silently disappear, the survivor keeps the combined photo provenance downstream, and review surfaces every merge with a veto that restores the row as a distinct approved match.
- **Version resolution**: once a game ID is settled (auto or approved), fetch `/thing?id=<id>&versions=1` and score the version list against the extraction's edition cues (publisher, year, language, edition wording). Same three-way classification: a single clear winner is `version_auto`; multiple plausible → `version_ambiguous` (goes to review); no cues at all → `version_unknown` (acceptable — BGG allows collection entries with no version set, and guessing wrong is worse than leaving it blank).
- Output: `matches.csv` with columns: `title_raw, bgg_id, bgg_name, year, type, match_status (auto|ambiguous|unmatched|approved|rejected), version_id, version_name, version_status (version_auto|version_ambiguous|version_unknown|version_approved), candidates_json, version_candidates_json, source_photos`.
### Stage 3 — `review`: Human review of ambiguous/unmatched
- A minimal local review flow. A TUI is fine (e.g., `rich`/`textual`), or a tiny localhost web page — builder's choice, but keep it dependency-light. *(As built: both — a rich TUI and a six-page FastAPI app sharing one decision engine.)*
- For each `ambiguous` item: show the raw title, source photo filename, and candidate list (name, year, type, BGG rank, owned count) → pick one, skip, or reject.
- For each `unmatched` item: allow manual BGG ID entry or a free-text re-search.
- For each `version_ambiguous` item: show the version candidates (version name, publisher, year, language) next to the edition cues from the photo → pick one, or mark `version_unknown`. Keep this pass optional/skippable — version review shouldn't block getting games uploaded.
- Show the source photo (or a crop) alongside if easy; otherwise the filename is acceptable.
- Decisions update `matches.csv` in place (`approved` with chosen `bgg_id`, or `rejected`).
- Must be resumable — quitting mid-review loses nothing.
### Stage 4 — `diff`: Compare against existing BGG collection
- Fetch my current collection: `https://boardgamegeek.com/xmlapi2/collection?username=<me>&own=1` (handle the 202-retry queue; also pass `&subtype=boardgameexpansion` in a second call — the collection endpoint excludes expansions from the default subtype).
- Output `to_add.csv`: approved/auto matches whose IDs are **not** already in the collection.
- **Multiple editions of the same game**: collection items are identified by `collid` (one per copy), not just `objectid`. If matches contain two entries for the same `bgg_id` with different `version_id`s, both belong in the collection as separate entries. Diff logic: a (bgg_id, version_id) pair is "already owned" only if a collection item matches both; a bare bgg_id with `version_unknown` is "already owned" if any copy of that game exists.
- **Improvement pass (to_update)**: for games already owned whose collection entry has NO version set, where photo matching produced a `version_auto`/`version_approved` — emit `to_update.csv` (`collid, bgg_id, bgg_name, version_id, version_name`). This upgrades the hand-entered 2018 entries with edition data from the shelves. Strictly additive: only fill empty version fields; if the collection entry already has a version, never touch it (even if the photo disagrees — report the disagreement in the summary instead).
- Informational only: list collection entries not seen in any photo (possible missing/loaned/sold games) in the summary. No action taken.
- Print a summary: N recognized, N already owned, N to add (including second editions), N version updates, N rejected/unmatched.
### Stage 5 — `upload`: Add games via Playwright
- Log in to boardgamegeek.com with credentials from env vars (`BGG_USERNAME`, `BGG_PASSWORD`). Never write credentials to disk or logs. Persist the browser session/storage state locally so repeat runs don't re-login.
- For each row in `to_add.csv`: navigate to the game page, use the "Add to Collection" flow, set status **Owned**, and — when a `version_id` is present — set the specific version in the collection item's version picker before saving. Manually walk this flow once and document the selectors before automating; the version UI is the most fragile part.
- Adding a second copy of an already-owned game must create a NEW collection entry, not edit the existing one.
- **Update mode** (rows from `to_update.csv`): open the EXISTING collection entry (keyed by `collid`) rather than the add flow, set the version, save. Must never create a duplicate entry and never change any other field of the entry. The verified route (2026-08-06, documented in `docs/bgg-upload-flow.md`): the collection table's version cell (`td.collection_version[onclick*="<collid>"]`) opens an inline editor whose radio values ARE version ids; clicking a radio saves immediately.
- Log every attempt to `upload_log.csv`: `bgg_id, name, status (added|already_present|failed), timestamp, error`.
- Idempotent: skip IDs already logged as `added`; re-verify against a fresh collection fetch on `--verify`.
- Deliberately slow: 24 s randomized delay between games. This is a real account on a community site — behave like a polite human.
- Expect UI fragility: fail gracefully per-game and continue; a `--retry-failed` flag re-attempts failures.
- Dry-run mode (`--dry-run`) that logs what it *would* add without touching the site.
### Stage 6 — `enrich`: Capture full game metadata
- For every approved/auto game ID, fetch `/thing?id=<batched,ids>&stats=1` (comma-separated batches of ~20) and store the critical data: name, year published, designers, artists, publishers, min/max players, community best-player-counts, playtime, min age, weight/complexity, BGG rank + rating, categories, mechanics, description, image + thumbnail URLs — plus the chosen version's details (version name, publisher, year, language) when known.
- Output: `data/games.json`, keyed by bgg_id (+ version_id where set). This file is the seed for the future web frontend, so keep it complete and stable.
- Idempotent and cheap: everything comes through the existing cache; `--refresh` forces a re-fetch (ranks and ratings drift over time).
## Data Model
All artifacts are flat files in a `data/` directory — human-readable, git-friendly, and reusable by the future frontend:
- `titles.json` — extraction output (stage 1)
- `unidentified.json` — game boxes seen but not identified (stage 1); retake prompts
- `title_edits.json`, `title_splits.json`, `title_removals.json` — durable human curation (corrected reads/cues, split-into-copies decisions, removed lines); replayed on every titles.json rebuild so re-extraction never undoes them
- `unidentified_dismissed.json` — dismissed retake prompts (kept apart so rebuilds can't resurrect them)
- `bgg_cache/` — cached XML API responses
- `matches.csv` — the master matching table (stages 23)
- `to_add.csv` — upload queue, new entries (stage 4)
- `to_update.csv` — version upgrades for existing version-less entries (stage 4)
- `upload_log.csv` — audit trail (stage 5)
- `games.json` — full game + version metadata (stage 6); seed data for the future frontend
## Configuration
`config.toml` (or env) for: BGG username, photo directory, model name, rate-limit settings. Secrets only via env vars.
## Error Handling & Edge Cases
- **202 queue** on `/collection`: retry with backoff (2s, 5s, 10s, 30s; give up after ~5 tries with a clear message).
- **HEIC photos** from iPhone: convert transparently.
- **Duplicate copies**: a title appearing in multiple photos with consistent edition cues is one game (dedupe). But genuinely distinct editions of the same game ARE in scope — they stay separate entries end-to-end and become separate BGG collection entries. *(Amended as built:)* identical duplicate copies of the same edition are supported too, as an explicit human decision — a photo-scoped "split into copies" recorded in `data/title_splits.json`, honored by every dedupe pass thereafter.
- **Non-game items** on shelves (books, card sleeves, storage boxes): the vision prompt should be instructed to include only board/card games; anything that slips through will fail resolution and land in review.
- **Special characters** in titles (é, colons, ampersands): normalize consistently on both sides of the match.
- **Base game vs. expansion vs. new edition**: the most common failure mode. Bias toward surfacing these as `ambiguous` rather than auto-matching.
## Acceptance Criteria
1. Given a directory of shelf photos, `bggpipe extract && bggpipe resolve` produces `matches.csv` with ≥90% of clearly legible titles auto-matched correctly.
2. Review flow lets me resolve every ambiguous/unmatched item without editing CSVs by hand.
3. `bggpipe upload --dry-run` shows exactly what would be added; the real run adds them, and a subsequent `bggpipe diff` reports zero remaining.
4. Killing any stage mid-run and restarting loses no work.
5. No BGG endpoint is hit faster than the rate limits above.
6. Credentials never appear in any file, log, or error message.
7. Where photos show legible edition cues, the matched version survives to the BGG collection entry; where they don't, the entry is added version-less rather than with a guessed version.
8. A game I own in two editions ends up as two distinct collection entries, and `games.json` contains full metadata for every game in the collection.
9. Version upgrades land on existing collection entries (same `collid`) with no duplicate entries created and no non-version fields changed; entries that already have a version are never modified.
## Suggested Build Order
1. Scaffold CLI + config + BGG API client (with caching, 202 handling, rate limiting). Test against my real username read-only.
2. Stage 2 resolve with a hand-typed test title list — validates matching before spending vision tokens.
3. Stage 1 extract against 23 test photos; iterate on the prompt.
4. Stage 3 review + Stage 4 diff.
5. Stage 5 upload — test with `--dry-run`, then a single game, then small batches. Verify the version-picker flow manually on one game first.
6. Stage 6 enrich — mostly free once the cached API client exists; build alongside stage 4.