Files
bggpipe/CLAUDE.md
T
Eric WagonerandClaude Fable 5 7e95ed607d RPGs pull real RPGGeek data; off-BGG games get facts and a cover photo
Two gaps at the edges of the library, both closed.

RPGGeek items live in the same database but use their own link types —
rpgdesigner, rpgpublisher, rpggenre, rpgcategory, rpgmechanic — so a
board-game-only parser found none of them and both RPG entries showed
just a year and a description. parse_things_full now reads both
vocabularies (plus rpgproducer/rpgseries): .dungeon gains John Battle
and Project Nerves, Parsely gains Jared A. Sorensen and its genres.

An off-BGG game has no API to enrich it and no publisher art to fetch,
so its detail page now hosts the only source it will ever have: a form
for title, year, players, playing time, publishers, designers and
notes, plus a cover photo upload. Both persist in data/local_games.json
and data/local_art/ (committed, like every other curation store) and
enrich merges them over the photo reads, so a rebuild can't erase them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jXZFSTZQKzAC8fqpWSz9g
2026-08-05 23:45:09 -04:00

7.7 KiB
Raw Blame History

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Project

bggpipe — a resumable, idempotent CLI pipeline that turns shelf photos into a BoardGameGeek collection, in six stages:

# Command What it does Status
1 bggpipe extract Claude vision reads titles + edition cues from photos/ working
2 bggpipe resolve match titles to BGG IDs and versions via XML API2 working (real API data since 2026-08-05)
3 bggpipe review human review of ambiguous/unmatched items; --web serves a FastAPI UI on port 8377 working
4 bggpipe diff diff approved matches against the existing BGG collection working
5 bggpipe upload add games via a logged-in Playwright session built; browser flows unverified until real data exists (--dry-run works now)
6 bggpipe enrich fetch full game/version metadata into games.json working

Full design lives in bgg-shelf-pipeline-spec.md (read it before changing pipeline semantics); the upload-stage walkthrough is in docs/bgg-upload-flow.md.

Commands

  • uv sync — install deps (Python 3.12+, managed by uv; use uv add, never pip). uv run bggpipe init handles first-run setup (folders, .env credentials, the one-time playwright install chromium).
  • uv run bggpipe web — the app: seven pages (Pipeline /, Photos, Titles, Review, Queue, Library, Help) in a shared sidebar shell (responsive: hamburger nav + stacked tables under 900px); stage runs execute one-at-a-time in a background job. --lan binds 0.0.0.0 behind a per-device access key: persisted in data/.lan_key (gitignored), printed as a QR at startup, cookie-paired for a year, required on EVERY network request (loopback clients and /static/* are exempt; the Host/Origin guard still applies). Phone camera uploads (generic image.jpg names) get minted shelf-<timestamp> names — only explicitly-named files trigger the replace-to-reshoot flow.
  • uv run bggpipe <stage> — run a pipeline stage. Non-secret settings come from config.toml (dirs, rate limit, vision_provider + per-provider [vision.*] blocks — "anthropic" or any OpenAI-compatible endpoint incl. local Ollama); --config overrides the path.
  • uv run pytest — the suite runs fully offline against fixtures. Tests marked live hit the real BGG API (read-only) and are skipped unless you pass --run-live.
  • uv run ruff check / uv run ruff format — lint (rules E, F, I, UP, B, SIM) and format.

Layout

  • src/bggpipe/cli.py (typer app), one module per stage (extract, resolve, review + webreview, diff, upload, enrich); templates/shell.html + templates/pages/* + static/app.{css,js} are the web UI (the stylesheet is the design system — tokens derive from the mascot art), plus bgg_client.py (rate-limited XML API2 client that caches responses to data/bgg_cache/), jobs.py (single-slot background stage runner for the web UI), normalize.py (title normalization), models.py (dataclasses), config.py, fsio.py (atomic writes), init_wizard.py (first-run setup).
  • scripts/write_stub_fixtures.py / write_photo_fixtures.py generate synthetic fixtures; record_fixtures.py re-records real API responses once a token exists.
  • tests/fixtures/bgg_cache/ — stub XML fixtures the offline tests run against.
  • data/ — pipeline state (CSV/JSON artifacts are committed; caches are not — see Git).

Hard rules (from spec — never violate)

  • ≤1 request every 2 seconds to any BGG endpoint; jittered backoff on 429/503. Upload stage: 24 s randomized delay between games.
  • Credentials never touch disk or logs. ANTHROPIC_API_KEY, BGG_USERNAME, BGG_PASSWORD, BGG_API_TOKEN come from env vars only. Playwright storage state is credential-adjacent — keep it gitignored.
  • The XML API requires a registered app token (Authorization: Bearer, from BGG_API_TOKEN) — unregistered requests get 401. Until Eric's registration at boardgamegeek.com/applications is approved, tests run on the stub fixtures in tests/fixtures/bgg_cache/; re-record them with scripts/record_fixtures.py once the token exists.
  • Every stage is idempotent and resumable — killing mid-run and restarting must lose no work; re-runs skip already-processed items.
  • Use only the XML API2 and the public website — no undocumented BGG endpoints (BGG tightened access policies in 2025).
  • BGG has no write API: writes drive the real website with a logged-in Playwright session.
  • Synthetic data must never reach upload. The stub era ended 2026-08-05: matches.csv and tests/fixtures/bgg_cache/ now hold real API data. The guard mechanism stays armed: the stub-fixture generators write data/bgg_cache/STUB_FIXTURES.marker (gitignored) and data/STUB_DATA.marker (committed), and the upload stage MUST refuse to run while either exists — regenerating stubs re-locks upload automatically.

Domain gotchas

  • Base game vs. expansion vs. new edition is the top failure mode — bias matching toward ambiguous over auto-match ("Wingspan Europe" must not match base Wingspan).
  • Editions/versions matter: Eric owns multiple editions of some games — each is a separate collection entry (keyed by collid on BGG). Never guess a version: no legible cues → version_unknown and a version-less collection entry.
  • Normalize titles (casefold, strip punctuation/articles, special chars like é/&/:) identically on both sides of a match; dedupe across photos but keep source_photos provenance.
  • Human curation is durable: data/title_splits.json (photo-scoped split-into-copies decisions, honored by extract's dedupe AND resolve's dedupe), data/title_edits.json (corrected reads/cues, applied before dedupe on every titles.json rebuild), data/title_removals.json (lines removed from the catalog — filtered out of every rebuild; delete the record to undo), data/title_additions.json (games added without a photo — joined into every rebuild; a later photo sighting dedupe-merges with them), and data/local_games.json + data/local_art/ (hand-written facts and a cover photo for off-BGG games — the ONLY source for them, merged over the photo reads by enrich) persist forever. Row-level decisions persist via the dedupe_veto column — edits never drop veto'd rows (a rename retitles them in place); removal drops them (explicitly discarding the line).
  • RPGs are local-only citizens: when the board-game search runs dry, resolve falls back to type=rpgitem (same geekdo API/token). RPGGeek items carry their OWN link types (rpgdesigner, rpgpublisher, rpggenre, rpgcategory, rpgmechanic) — a board-game-only parser silently returns nothing for them. Matched rpgitems enrich into the library but diff routes them to local_only — they must never reach to_add.csv/upload (their collection lives on RPGGeek, out of scope).
  • Detailed BGG API behavior (202 queueing, collection-endpoint quirks, endpoints): use the bgg-api skill. If the spec's BGG behavior changes, update the bgg-api skill to match — they must not drift.

Git

  • Remote is self-hosted Gitea 1.26 (git.kestrelsnest.social/eric/bggpipe), not GitHubgh CLI does not work here.
  • Commit data/matches.csv, data/to_add.csv, data/to_update.csv, data/upload_log.csv, data/titles.json, data/unidentified.json, data/unidentified_dismissed.json, data/title_splits.json, data/title_edits.json, data/title_removals.json, data/title_additions.json, data/local_games.json, data/local_art/, data/games.json, data/STUB_DATA.marker (while it applies), and the collection snapshot XMLs. Never commit data/bgg_cache/, data/extract_raw/, photos/, data/.lan_key, Playwright storage state, or .env.