Files
bggpipe/CLAUDE.md
T
Eric Wagoner 74fc4fe847 RPGs become local library citizens: identified, enriched, never uploaded
When the board-game search (and truncation heads) runs dry, resolve
falls back to type=rpgitem — the geekdo database is shared, so the same
API, token, cache, and classification machinery apply. Matched rpgitems
flow through review and enrich normally but diff routes them to a
local_only bucket, structurally outside to_add/to_update: their
collections live on RPGGeek, beyond this pipeline's write scope. The
library page gains an All/Board games/RPGs filter and an "RPG · local
only" badge; the catalog tags them too. Fixture generators write blanket
empty rpgitem stubs for every known query (the fallback fires for every
unmatched title) with real synthetic entries for Alice Is Missing.

Data: both Alice rows re-resolved from unmatched to auto rpgitem
matches. First diff since the audit reworks also lands their real-data
consequences: Dungeon! gains its TSR edition update on a versionless
copy the old claim ordering missed, to_add rows carry unioned reshoot
provenance, and the Herbaceous typo row's survivor is now the
correctly-spelled title.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 19:01:09 -04:00

6.2 KiB
Raw Blame History

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Project

bggpipe — a resumable, idempotent CLI pipeline that turns shelf photos into a BoardGameGeek collection, in six stages:

# Command What it does Status
1 bggpipe extract Claude vision reads titles + edition cues from photos/ working
2 bggpipe resolve match titles to BGG IDs and versions via XML API2 working (stub data — see hard rules)
3 bggpipe review human review of ambiguous/unmatched items; --web serves a FastAPI UI on port 8377 working
4 bggpipe diff diff approved matches against the existing BGG collection working
5 bggpipe upload add games via a logged-in Playwright session built; browser flows unverified until real data exists (--dry-run works now)
6 bggpipe enrich fetch full game/version metadata into games.json working

Full design lives in bgg-shelf-pipeline-spec.md (read it before changing pipeline semantics); the upload-stage walkthrough is in docs/bgg-upload-flow.md.

Commands

  • uv sync — install deps (Python 3.12+, managed by uv; use uv add, never pip). uv run bggpipe init handles first-run setup (folders, .env credentials, the one-time playwright install chromium).
  • uv run bggpipe web — the app: six pages (Pipeline /, Photos, Review, Catalog, Queue, Library) in a shared sidebar shell; stage runs execute one-at-a-time in a background job.
  • uv run bggpipe <stage> — run a pipeline stage. Non-secret settings come from config.toml (username, dirs, vision model, rate limit); --config overrides the path.
  • uv run pytest — the suite runs fully offline against fixtures. Tests marked live hit the real BGG API (read-only) and are skipped unless you pass --run-live.
  • uv run ruff check / uv run ruff format — lint (rules E, F, I, UP, B, SIM) and format.

Layout

  • src/bggpipe/cli.py (typer app), one module per stage (extract, resolve, review + webreview, diff, upload, enrich); templates/shell.html + templates/pages/* + static/app.{css,js} are the web UI (the stylesheet is the design system — tokens derive from the mascot art), plus bgg_client.py (rate-limited XML API2 client that caches responses to data/bgg_cache/), jobs.py (single-slot background stage runner for the web UI), normalize.py (title normalization), models.py (dataclasses), config.py, fsio.py (atomic writes), init_wizard.py (first-run setup).
  • scripts/write_stub_fixtures.py / write_photo_fixtures.py generate synthetic fixtures; record_fixtures.py re-records real API responses once a token exists.
  • tests/fixtures/bgg_cache/ — stub XML fixtures the offline tests run against.
  • data/ — pipeline state (CSV/JSON artifacts are committed; caches are not — see Git).

Hard rules (from spec — never violate)

  • ≤1 request every 2 seconds to any BGG endpoint; jittered backoff on 429/503. Upload stage: 24 s randomized delay between games.
  • Credentials never touch disk or logs. ANTHROPIC_API_KEY, BGG_USERNAME, BGG_PASSWORD, BGG_API_TOKEN come from env vars only. Playwright storage state is credential-adjacent — keep it gitignored.
  • The XML API requires a registered app token (Authorization: Bearer, from BGG_API_TOKEN) — unregistered requests get 401. Until Eric's registration at boardgamegeek.com/applications is approved, tests run on the stub fixtures in tests/fixtures/bgg_cache/; re-record them with scripts/record_fixtures.py once the token exists.
  • Every stage is idempotent and resumable — killing mid-run and restarting must lose no work; re-runs skip already-processed items.
  • Use only the XML API2 and the public website — no undocumented BGG endpoints (BGG tightened access policies in 2025).
  • BGG has no write API: writes drive the real website with a logged-in Playwright session.
  • Stub-resolved data is never upload-ready. All version_ids (and some game data) in matches.csv, to_add.csv, and to_update.csv currently come from SYNTHETIC stub fixtures — placeholders until real fixtures exist. When BGG_API_TOKEN arrives: delete both cache dirs, re-record fixtures, resolve --force, re-review. Two provenance markers guard this (both written by the fixture generators): data/bgg_cache/STUB_FIXTURES.marker (gitignored, travels with the stub XML) and data/STUB_DATA.marker (committed, so a fresh clone stays guarded). The upload stage MUST refuse to run while either exists; delete data/STUB_DATA.marker only after re-resolving from real fixtures.

Domain gotchas

  • Base game vs. expansion vs. new edition is the top failure mode — bias matching toward ambiguous over auto-match ("Wingspan Europe" must not match base Wingspan).
  • Editions/versions matter: Eric owns multiple editions of some games — each is a separate collection entry (keyed by collid on BGG). Never guess a version: no legible cues → version_unknown and a version-less collection entry.
  • Normalize titles (casefold, strip punctuation/articles, special chars like é/&/:) identically on both sides of a match; dedupe across photos but keep source_photos provenance.
  • RPGs are local-only citizens: when the board-game search runs dry, resolve falls back to type=rpgitem (same geekdo API/token). Matched rpgitems enrich into the library but diff routes them to local_only — they must never reach to_add.csv/upload (their collection lives on RPGGeek, out of scope).
  • Detailed BGG API behavior (202 queueing, collection-endpoint quirks, endpoints): use the bgg-api skill. If the spec's BGG behavior changes, update the bgg-api skill to match — they must not drift.

Git

  • Remote is self-hosted Gitea 1.26 (git.kestrelsnest.social/eric/bggpipe), not GitHubgh CLI does not work here.
  • Commit data/matches.csv, data/to_add.csv, data/to_update.csv, data/upload_log.csv, data/titles.json, data/unidentified.json, data/unidentified_dismissed.json, data/games.json, data/STUB_DATA.marker (while it applies), and the collection snapshot XMLs. Never commit data/bgg_cache/, data/extract_raw/, photos/, Playwright storage state, or .env.