Files
bggpipe/bgg-shelf-pipeline-spec.md
T
Eric Wagoner 8420a0a1ca Post-resolve dedupe: duplicate reads merge, review can veto
Rows resolving to the same (bgg_id, version_id — or both version-
unknown) are the same physical game read twice unless their extraction
cues conflict (two editions stay separate). The survivor is the read
whose transcription matches the BGG name; losers are marked
match_status=merged with a new merged_into column — no row is ever
deleted, and older matches.csv files without the column still read.
Downstream: diff skips merged rows but folds their photos into the
survivor's to_add provenance; enrich and the review passes ignore them.
The web UI gains a Merges section ("Jokin Ha... merged into Joking
Hazard") with a veto (v key) that restores the row as a distinct
approved match, plus a merged catalog chip and header tally.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 16:59:47 -04:00

16 KiB
Raw Blame History

Shelf-to-BGG Collection Pipeline — Build Spec

Goal

Build a command-line pipeline that takes photos of my board game shelves and ends with every recognized game added to my BoardGameGeek collection. The pipeline must be resumable, idempotent, and leave a human-reviewable audit trail at every stage.

Beyond the bare game, capture which edition/version I own wherever the photos allow it — many of my games exist in multiple editions, and in some cases I own more than one edition of the same game (each is a separate collection entry). Also capture all critical data about each game itself (see Stage 6 — enrich): the BGG /thing metadata is cheap to fetch and will seed the future frontend. Purchase provenance (where/when acquired, price paid) is explicitly NOT tracked — I don't have that data.

Out of scope (for now): the web frontend for displaying the collection. That's a follow-up project. But keep the data artifacts (see Data Model) clean and structured so they can seed it later.

Constraints & Context

  • BGG has no write API. Reads go through the XML API2 (https://boardgamegeek.com/xmlapi2/); writes must automate the website itself with a real login session.
  • The XML API requires a registered application (boardgamegeek.com/using_the_xml_api, policy 2025-07-02): every request carries Authorization: Bearer <token> from the BGG_API_TOKEN env var, sent to boardgamegeek.com without a leading www. Register a free non-commercial application at boardgamegeek.com/applications — approval can take a week+, so development runs on recorded/stub XML fixtures until then.
  • BGG's XML API queues collection requests: a first call may return HTTP 202 ("try again"). Retry with backoff.
  • BGG will throttle aggressive clients. Target ≤1 request every 2 seconds to any BGG endpoint, with jittered backoff on 429/503.
  • BGG changed API access policies in 2025; some older community tools broke. Don't depend on undocumented endpoints beyond XML API2 and the public website.
  • Vision extraction uses the Anthropic API (Claude with vision). Assume ANTHROPIC_API_KEY in the environment.
  • Runs on macOS. Prefer Python 3.12+ with uv for dependency management. Browser automation via Playwright (not Selenium).

Pipeline Overview

Five stages, each a separate subcommand of one CLI (suggest bggpipe):

photos/ → [1 extract] → titles.json → [2 resolve] → matches.csv
        → [3 review] → matches.csv (approved) → [4 diff] → to_add.csv
        → [5 upload] → upload_log.csv
        → [6 enrich] → games.json   (runs any time after review)

Each stage reads the previous stage's artifact and writes its own. Re-running a stage must be safe (skip already-processed items).

Stage 1 — extract: Vision title extraction

  • Input: a directory of shelf photos (JPEG/PNG/HEIC — convert HEIC to JPEG first via sips or Pillow-HEIF).
  • For each photo, send to the Anthropic API (model: latest Sonnet) with a prompt that asks for:
    • Every board game title visible (spines and face-out boxes), transcribed as printed.
    • A confidence level per title (high / medium / low).
    • Edition/version cues, each if legible: publisher name or logo, edition wording ("2nd Edition", "Deluxe", "Big Box", anniversary marks), copyright/print year, language, and distinctive box-art notes (colorway, artwork style). These drive version matching in Stage 2.
    • If the SAME title appears more than once in the photos with visibly different boxes, report each as a separate entry — I own multiple editions of some games.
  • Downscale images so the long edge is ≤1568px before sending (API sweet spot; keeps tokens down).
  • Prompt for structured JSON output; parse defensively (strip code fences).
  • Handle overlap: the same game may appear in two photos. Dedupe by normalized title (casefold, strip punctuation/articles) but keep provenance — record which photo(s) each title came from.
  • Output: titles.json — list of {title_raw, title_normalized, confidence, publisher_hint, edition_hint, year_hint, language_hint, art_notes, source_photos[]}.
  • Dedupe nuance: identical normalized titles from different photos collapse to one entry ONLY if their edition cues don't conflict; conflicting cues (different publisher/edition text) stay as separate entries.
  • Support --only <photo> to re-run a single photo (e.g., after retaking a blurry shot).
  • Unidentified sightings: boxes that appear to be games but can't be confidently titled (blurry, obscured, sharp angle, frame edge) are reported rather than silently omitted — location described relative to identified neighbors, plus any partial text and art notes → unidentified.json, keyed by photo. The end-of-run summary lists them (and low-confidence reads) so I can take a closer photo and re-run with --only.

Stage 2 — resolve: Match titles to BGG IDs

  • For each title, query https://boardgamegeek.com/xmlapi2/search?query=<title>&type=boardgame,boardgameexpansion.
  • Scoring heuristic for candidates:
    1. Exact normalized-name match → strong.
    2. Fuzzy match (e.g., rapidfuzz token_sort_ratio ≥ 90) → good.
    3. If multiple candidates tie, fetch /thing?id=...&stats=1 for the top ~5 and prefer higher-owned/higher-ranked entries (obscure duplicates lose to the well-known game of the same name).
  • Classify each result:
    • auto — single confident match, no review needed.
    • ambiguous — multiple plausible matches; store all candidates.
    • unmatched — nothing plausible found.
  • Expansions: BGG returns boardgameexpansion as a distinct type. Keep them — I own expansions and want them in the collection — but tag them so review can catch base-game/expansion confusion (a spine reading "Wingspan Europe" must not match base Wingspan).
  • Cache all BGG responses on disk (keyed by query/ID) so re-runs don't re-hit the API.
  • Post-resolve dedupe: rows resolving to the same (bgg_id, version_id — or both version-unknown) are the same physical game read twice (typo, partial spine) unless their extraction cues conflict (two editions). Losers get match_status=merged + a merged_into column pointing at the survivor — rows never silently disappear, the survivor keeps the combined photo provenance downstream, and review surfaces every merge with a veto that restores the row as a distinct approved match.
  • Version resolution: once a game ID is settled (auto or approved), fetch /thing?id=<id>&versions=1 and score the version list against the extraction's edition cues (publisher, year, language, edition wording). Same three-way classification: a single clear winner is version_auto; multiple plausible → version_ambiguous (goes to review); no cues at all → version_unknown (acceptable — BGG allows collection entries with no version set, and guessing wrong is worse than leaving it blank).
  • Output: matches.csv with columns: title_raw, bgg_id, bgg_name, year, type, match_status (auto|ambiguous|unmatched|approved|rejected), version_id, version_name, version_status (version_auto|version_ambiguous|version_unknown|version_approved), candidates_json, version_candidates_json, source_photos.

Stage 3 — review: Human review of ambiguous/unmatched

  • A minimal local review flow. A TUI is fine (e.g., rich/textual), or a tiny localhost web page — builder's choice, but keep it dependency-light.
  • For each ambiguous item: show the raw title, source photo filename, and candidate list (name, year, type, BGG rank, owned count) → pick one, skip, or reject.
  • For each unmatched item: allow manual BGG ID entry or a free-text re-search.
  • For each version_ambiguous item: show the version candidates (version name, publisher, year, language) next to the edition cues from the photo → pick one, or mark version_unknown. Keep this pass optional/skippable — version review shouldn't block getting games uploaded.
  • Show the source photo (or a crop) alongside if easy; otherwise the filename is acceptable.
  • Decisions update matches.csv in place (approved with chosen bgg_id, or rejected).
  • Must be resumable — quitting mid-review loses nothing.

Stage 4 — diff: Compare against existing BGG collection

  • Fetch my current collection: https://boardgamegeek.com/xmlapi2/collection?username=<me>&own=1 (handle the 202-retry queue; also pass &subtype=boardgameexpansion in a second call — the collection endpoint excludes expansions from the default subtype).
  • Output to_add.csv: approved/auto matches whose IDs are not already in the collection.
  • Multiple editions of the same game: collection items are identified by collid (one per copy), not just objectid. If matches contain two entries for the same bgg_id with different version_ids, both belong in the collection as separate entries. Diff logic: a (bgg_id, version_id) pair is "already owned" only if a collection item matches both; a bare bgg_id with version_unknown is "already owned" if any copy of that game exists.
  • Improvement pass (to_update): for games already owned whose collection entry has NO version set, where photo matching produced a version_auto/version_approved — emit to_update.csv (collid, bgg_id, bgg_name, version_id, version_name). This upgrades the hand-entered 2018 entries with edition data from the shelves. Strictly additive: only fill empty version fields; if the collection entry already has a version, never touch it (even if the photo disagrees — report the disagreement in the summary instead).
  • Informational only: list collection entries not seen in any photo (possible missing/loaned/sold games) in the summary. No action taken.
  • Print a summary: N recognized, N already owned, N to add (including second editions), N version updates, N rejected/unmatched.

Stage 5 — upload: Add games via Playwright

  • Log in to boardgamegeek.com with credentials from env vars (BGG_USERNAME, BGG_PASSWORD). Never write credentials to disk or logs. Persist the browser session/storage state locally so repeat runs don't re-login.
  • For each row in to_add.csv: navigate to the game page, use the "Add to Collection" flow, set status Owned, and — when a version_id is present — set the specific version in the collection item's version picker before saving. Manually walk this flow once and document the selectors before automating; the version UI is the most fragile part.
  • Adding a second copy of an already-owned game must create a NEW collection entry, not edit the existing one.
  • Update mode (rows from to_update.csv): open the EXISTING collection entry (keyed by collid) rather than the add flow, set the version, save. Must never create a duplicate entry and never change any other field of the entry. Verify the already-owned dialog behavior manually first — flagged as unverified in docs/bgg-upload-flow.md.
  • Log every attempt to upload_log.csv: bgg_id, name, status (added|already_present|failed), timestamp, error.
  • Idempotent: skip IDs already logged as added; re-verify against a fresh collection fetch on --verify.
  • Deliberately slow: 24 s randomized delay between games. This is a real account on a community site — behave like a polite human.
  • Expect UI fragility: fail gracefully per-game and continue; a --retry-failed flag re-attempts failures.
  • Dry-run mode (--dry-run) that logs what it would add without touching the site.

Stage 6 — enrich: Capture full game metadata

  • For every approved/auto game ID, fetch /thing?id=<batched,ids>&stats=1 (comma-separated batches of ~20) and store the critical data: name, year published, designers, artists, publishers, min/max players, community best-player-counts, playtime, min age, weight/complexity, BGG rank + rating, categories, mechanics, description, image + thumbnail URLs — plus the chosen version's details (version name, publisher, year, language) when known.
  • Output: data/games.json, keyed by bgg_id (+ version_id where set). This file is the seed for the future web frontend, so keep it complete and stable.
  • Idempotent and cheap: everything comes through the existing cache; --refresh forces a re-fetch (ranks and ratings drift over time).

Data Model

All artifacts are flat files in a data/ directory — human-readable, git-friendly, and reusable by the future frontend:

  • titles.json — extraction output (stage 1)
  • unidentified.json — game boxes seen but not identified (stage 1); retake prompts
  • bgg_cache/ — cached XML API responses
  • matches.csv — the master matching table (stages 23)
  • to_add.csv — upload queue, new entries (stage 4)
  • to_update.csv — version upgrades for existing version-less entries (stage 4)
  • upload_log.csv — audit trail (stage 5)
  • games.json — full game + version metadata (stage 6); seed data for the future frontend

Configuration

config.toml (or env) for: BGG username, photo directory, model name, rate-limit settings. Secrets only via env vars.

Error Handling & Edge Cases

  • 202 queue on /collection: retry with backoff (2s, 5s, 10s, 30s; give up after ~5 tries with a clear message).
  • HEIC photos from iPhone: convert transparently.
  • Duplicate copies: a title appearing in multiple photos with consistent edition cues is one game (dedupe). But genuinely distinct editions of the same game ARE in scope — they stay separate entries end-to-end and become separate BGG collection entries. Identical duplicate copies of the same edition are out of scope (assume dedupe).
  • Non-game items on shelves (books, card sleeves, storage boxes): the vision prompt should be instructed to include only board/card games; anything that slips through will fail resolution and land in review.
  • Special characters in titles (é, colons, ampersands): normalize consistently on both sides of the match.
  • Base game vs. expansion vs. new edition: the most common failure mode. Bias toward surfacing these as ambiguous rather than auto-matching.

Acceptance Criteria

  1. Given a directory of shelf photos, bggpipe extract && bggpipe resolve produces matches.csv with ≥90% of clearly legible titles auto-matched correctly.
  2. Review flow lets me resolve every ambiguous/unmatched item without editing CSVs by hand.
  3. bggpipe upload --dry-run shows exactly what would be added; the real run adds them, and a subsequent bggpipe diff reports zero remaining.
  4. Killing any stage mid-run and restarting loses no work.
  5. No BGG endpoint is hit faster than the rate limits above.
  6. Credentials never appear in any file, log, or error message.
  7. Where photos show legible edition cues, the matched version survives to the BGG collection entry; where they don't, the entry is added version-less rather than with a guessed version.
  8. A game I own in two editions ends up as two distinct collection entries, and games.json contains full metadata for every game in the collection.
  9. Version upgrades land on existing collection entries (same collid) with no duplicate entries created and no non-version fields changed; entries that already have a version are never modified.

Suggested Build Order

  1. Scaffold CLI + config + BGG API client (with caching, 202 handling, rate limiting). Test against my real username read-only.
  2. Stage 2 resolve with a hand-typed test title list — validates matching before spending vision tokens.
  3. Stage 1 extract against 23 test photos; iterate on the prompt.
  4. Stage 3 review + Stage 4 diff.
  5. Stage 5 upload — test with --dry-run, then a single game, then small batches. Verify the version-picker flow manually on one game first.
  6. Stage 6 enrich — mostly free once the cached API client exists; build alongside stage 4.