Rows resolving to the same (bgg_id, version_id — or both version-
unknown) are the same physical game read twice unless their extraction
cues conflict (two editions stay separate). The survivor is the read
whose transcription matches the BGG name; losers are marked
match_status=merged with a new merged_into column — no row is ever
deleted, and older matches.csv files without the column still read.
Downstream: diff skips merged rows but folds their photos into the
survivor's to_add provenance; enrich and the review passes ignore them.
The web UI gains a Merges section ("Jokin Ha... merged into Joking
Hazard") with a veto (v key) that restores the row as a distinct
approved match, plus a merged catalog chip and header tally.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
16 KiB
Shelf-to-BGG Collection Pipeline — Build Spec
Goal
Build a command-line pipeline that takes photos of my board game shelves and ends with every recognized game added to my BoardGameGeek collection. The pipeline must be resumable, idempotent, and leave a human-reviewable audit trail at every stage.
Beyond the bare game, capture which edition/version I own wherever the photos allow it — many of my games exist in multiple editions, and in some cases I own more than one edition of the same game (each is a separate collection entry). Also capture all critical data about each game itself (see Stage 6 — enrich): the BGG /thing metadata is cheap to fetch and will seed the future frontend. Purchase provenance (where/when acquired, price paid) is explicitly NOT tracked — I don't have that data.
Out of scope (for now): the web frontend for displaying the collection. That's a follow-up project. But keep the data artifacts (see Data Model) clean and structured so they can seed it later.
Constraints & Context
- BGG has no write API. Reads go through the XML API2 (
https://boardgamegeek.com/xmlapi2/); writes must automate the website itself with a real login session. - The XML API requires a registered application (boardgamegeek.com/using_the_xml_api, policy 2025-07-02): every request carries
Authorization: Bearer <token>from theBGG_API_TOKENenv var, sent toboardgamegeek.comwithout a leadingwww. Register a free non-commercial application at boardgamegeek.com/applications — approval can take a week+, so development runs on recorded/stub XML fixtures until then. - BGG's XML API queues collection requests: a first call may return HTTP 202 ("try again"). Retry with backoff.
- BGG will throttle aggressive clients. Target ≤1 request every 2 seconds to any BGG endpoint, with jittered backoff on 429/503.
- BGG changed API access policies in 2025; some older community tools broke. Don't depend on undocumented endpoints beyond XML API2 and the public website.
- Vision extraction uses the Anthropic API (Claude with vision). Assume
ANTHROPIC_API_KEYin the environment. - Runs on macOS. Prefer Python 3.12+ with
uvfor dependency management. Browser automation via Playwright (not Selenium).
Pipeline Overview
Five stages, each a separate subcommand of one CLI (suggest bggpipe):
photos/ → [1 extract] → titles.json → [2 resolve] → matches.csv
→ [3 review] → matches.csv (approved) → [4 diff] → to_add.csv
→ [5 upload] → upload_log.csv
→ [6 enrich] → games.json (runs any time after review)
Each stage reads the previous stage's artifact and writes its own. Re-running a stage must be safe (skip already-processed items).
Stage 1 — extract: Vision title extraction
- Input: a directory of shelf photos (JPEG/PNG/HEIC — convert HEIC to JPEG first via
sipsor Pillow-HEIF). - For each photo, send to the Anthropic API (model: latest Sonnet) with a prompt that asks for:
- Every board game title visible (spines and face-out boxes), transcribed as printed.
- A confidence level per title (
high/medium/low). - Edition/version cues, each if legible: publisher name or logo, edition wording ("2nd Edition", "Deluxe", "Big Box", anniversary marks), copyright/print year, language, and distinctive box-art notes (colorway, artwork style). These drive version matching in Stage 2.
- If the SAME title appears more than once in the photos with visibly different boxes, report each as a separate entry — I own multiple editions of some games.
- Downscale images so the long edge is ≤1568px before sending (API sweet spot; keeps tokens down).
- Prompt for structured JSON output; parse defensively (strip code fences).
- Handle overlap: the same game may appear in two photos. Dedupe by normalized title (casefold, strip punctuation/articles) but keep provenance — record which photo(s) each title came from.
- Output:
titles.json— list of{title_raw, title_normalized, confidence, publisher_hint, edition_hint, year_hint, language_hint, art_notes, source_photos[]}. - Dedupe nuance: identical normalized titles from different photos collapse to one entry ONLY if their edition cues don't conflict; conflicting cues (different publisher/edition text) stay as separate entries.
- Support
--only <photo>to re-run a single photo (e.g., after retaking a blurry shot). - Unidentified sightings: boxes that appear to be games but can't be confidently titled (blurry, obscured, sharp angle, frame edge) are reported rather than silently omitted — location described relative to identified neighbors, plus any partial text and art notes →
unidentified.json, keyed by photo. The end-of-run summary lists them (and low-confidence reads) so I can take a closer photo and re-run with--only.
Stage 2 — resolve: Match titles to BGG IDs
- For each title, query
https://boardgamegeek.com/xmlapi2/search?query=<title>&type=boardgame,boardgameexpansion. - Scoring heuristic for candidates:
- Exact normalized-name match → strong.
- Fuzzy match (e.g.,
rapidfuzztoken_sort_ratio ≥ 90) → good. - If multiple candidates tie, fetch
/thing?id=...&stats=1for the top ~5 and prefer higher-owned/higher-ranked entries (obscure duplicates lose to the well-known game of the same name).
- Classify each result:
auto— single confident match, no review needed.ambiguous— multiple plausible matches; store all candidates.unmatched— nothing plausible found.
- Expansions: BGG returns
boardgameexpansionas a distinct type. Keep them — I own expansions and want them in the collection — but tag them so review can catch base-game/expansion confusion (a spine reading "Wingspan Europe" must not match base Wingspan). - Cache all BGG responses on disk (keyed by query/ID) so re-runs don't re-hit the API.
- Post-resolve dedupe: rows resolving to the same (bgg_id, version_id — or both version-unknown) are the same physical game read twice (typo, partial spine) unless their extraction cues conflict (two editions). Losers get
match_status=merged+ amerged_intocolumn pointing at the survivor — rows never silently disappear, the survivor keeps the combined photo provenance downstream, and review surfaces every merge with a veto that restores the row as a distinct approved match. - Version resolution: once a game ID is settled (auto or approved), fetch
/thing?id=<id>&versions=1and score the version list against the extraction's edition cues (publisher, year, language, edition wording). Same three-way classification: a single clear winner isversion_auto; multiple plausible →version_ambiguous(goes to review); no cues at all →version_unknown(acceptable — BGG allows collection entries with no version set, and guessing wrong is worse than leaving it blank). - Output:
matches.csvwith columns:title_raw, bgg_id, bgg_name, year, type, match_status (auto|ambiguous|unmatched|approved|rejected), version_id, version_name, version_status (version_auto|version_ambiguous|version_unknown|version_approved), candidates_json, version_candidates_json, source_photos.
Stage 3 — review: Human review of ambiguous/unmatched
- A minimal local review flow. A TUI is fine (e.g.,
rich/textual), or a tiny localhost web page — builder's choice, but keep it dependency-light. - For each
ambiguousitem: show the raw title, source photo filename, and candidate list (name, year, type, BGG rank, owned count) → pick one, skip, or reject. - For each
unmatcheditem: allow manual BGG ID entry or a free-text re-search. - For each
version_ambiguousitem: show the version candidates (version name, publisher, year, language) next to the edition cues from the photo → pick one, or markversion_unknown. Keep this pass optional/skippable — version review shouldn't block getting games uploaded. - Show the source photo (or a crop) alongside if easy; otherwise the filename is acceptable.
- Decisions update
matches.csvin place (approvedwith chosenbgg_id, orrejected). - Must be resumable — quitting mid-review loses nothing.
Stage 4 — diff: Compare against existing BGG collection
- Fetch my current collection:
https://boardgamegeek.com/xmlapi2/collection?username=<me>&own=1(handle the 202-retry queue; also pass&subtype=boardgameexpansionin a second call — the collection endpoint excludes expansions from the default subtype). - Output
to_add.csv: approved/auto matches whose IDs are not already in the collection. - Multiple editions of the same game: collection items are identified by
collid(one per copy), not justobjectid. If matches contain two entries for the samebgg_idwith differentversion_ids, both belong in the collection as separate entries. Diff logic: a (bgg_id, version_id) pair is "already owned" only if a collection item matches both; a bare bgg_id withversion_unknownis "already owned" if any copy of that game exists. - Improvement pass (to_update): for games already owned whose collection entry has NO version set, where photo matching produced a
version_auto/version_approved— emitto_update.csv(collid, bgg_id, bgg_name, version_id, version_name). This upgrades the hand-entered 2018 entries with edition data from the shelves. Strictly additive: only fill empty version fields; if the collection entry already has a version, never touch it (even if the photo disagrees — report the disagreement in the summary instead). - Informational only: list collection entries not seen in any photo (possible missing/loaned/sold games) in the summary. No action taken.
- Print a summary: N recognized, N already owned, N to add (including second editions), N version updates, N rejected/unmatched.
Stage 5 — upload: Add games via Playwright
- Log in to boardgamegeek.com with credentials from env vars (
BGG_USERNAME,BGG_PASSWORD). Never write credentials to disk or logs. Persist the browser session/storage state locally so repeat runs don't re-login. - For each row in
to_add.csv: navigate to the game page, use the "Add to Collection" flow, set status Owned, and — when aversion_idis present — set the specific version in the collection item's version picker before saving. Manually walk this flow once and document the selectors before automating; the version UI is the most fragile part. - Adding a second copy of an already-owned game must create a NEW collection entry, not edit the existing one.
- Update mode (rows from
to_update.csv): open the EXISTING collection entry (keyed bycollid) rather than the add flow, set the version, save. Must never create a duplicate entry and never change any other field of the entry. Verify the already-owned dialog behavior manually first — flagged as unverified indocs/bgg-upload-flow.md. - Log every attempt to
upload_log.csv:bgg_id, name, status (added|already_present|failed), timestamp, error. - Idempotent: skip IDs already logged as
added; re-verify against a fresh collection fetch on--verify. - Deliberately slow: 2–4 s randomized delay between games. This is a real account on a community site — behave like a polite human.
- Expect UI fragility: fail gracefully per-game and continue; a
--retry-failedflag re-attempts failures. - Dry-run mode (
--dry-run) that logs what it would add without touching the site.
Stage 6 — enrich: Capture full game metadata
- For every approved/auto game ID, fetch
/thing?id=<batched,ids>&stats=1(comma-separated batches of ~20) and store the critical data: name, year published, designers, artists, publishers, min/max players, community best-player-counts, playtime, min age, weight/complexity, BGG rank + rating, categories, mechanics, description, image + thumbnail URLs — plus the chosen version's details (version name, publisher, year, language) when known. - Output:
data/games.json, keyed by bgg_id (+ version_id where set). This file is the seed for the future web frontend, so keep it complete and stable. - Idempotent and cheap: everything comes through the existing cache;
--refreshforces a re-fetch (ranks and ratings drift over time).
Data Model
All artifacts are flat files in a data/ directory — human-readable, git-friendly, and reusable by the future frontend:
titles.json— extraction output (stage 1)unidentified.json— game boxes seen but not identified (stage 1); retake promptsbgg_cache/— cached XML API responsesmatches.csv— the master matching table (stages 2–3)to_add.csv— upload queue, new entries (stage 4)to_update.csv— version upgrades for existing version-less entries (stage 4)upload_log.csv— audit trail (stage 5)games.json— full game + version metadata (stage 6); seed data for the future frontend
Configuration
config.toml (or env) for: BGG username, photo directory, model name, rate-limit settings. Secrets only via env vars.
Error Handling & Edge Cases
- 202 queue on
/collection: retry with backoff (2s, 5s, 10s, 30s; give up after ~5 tries with a clear message). - HEIC photos from iPhone: convert transparently.
- Duplicate copies: a title appearing in multiple photos with consistent edition cues is one game (dedupe). But genuinely distinct editions of the same game ARE in scope — they stay separate entries end-to-end and become separate BGG collection entries. Identical duplicate copies of the same edition are out of scope (assume dedupe).
- Non-game items on shelves (books, card sleeves, storage boxes): the vision prompt should be instructed to include only board/card games; anything that slips through will fail resolution and land in review.
- Special characters in titles (é, colons, ampersands): normalize consistently on both sides of the match.
- Base game vs. expansion vs. new edition: the most common failure mode. Bias toward surfacing these as
ambiguousrather than auto-matching.
Acceptance Criteria
- Given a directory of shelf photos,
bggpipe extract && bggpipe resolveproducesmatches.csvwith ≥90% of clearly legible titles auto-matched correctly. - Review flow lets me resolve every ambiguous/unmatched item without editing CSVs by hand.
bggpipe upload --dry-runshows exactly what would be added; the real run adds them, and a subsequentbggpipe diffreports zero remaining.- Killing any stage mid-run and restarting loses no work.
- No BGG endpoint is hit faster than the rate limits above.
- Credentials never appear in any file, log, or error message.
- Where photos show legible edition cues, the matched version survives to the BGG collection entry; where they don't, the entry is added version-less rather than with a guessed version.
- A game I own in two editions ends up as two distinct collection entries, and
games.jsoncontains full metadata for every game in the collection. - Version upgrades land on existing collection entries (same
collid) with no duplicate entries created and no non-version fields changed; entries that already have a version are never modified.
Suggested Build Order
- Scaffold CLI + config + BGG API client (with caching, 202 handling, rate limiting). Test against my real username read-only.
- Stage 2 resolve with a hand-typed test title list — validates matching before spending vision tokens.
- Stage 1 extract against 2–3 test photos; iterate on the prompt.
- Stage 3 review + Stage 4 diff.
- Stage 5 upload — test with
--dry-run, then a single game, then small batches. Verify the version-picker flow manually on one game first. - Stage 6 enrich — mostly free once the cached API client exists; build alongside stage 4.