Files
bggpipe/docs/guide.md
T
Eric WagonerandClaude Fable 5 1d9baff989 Requirements answer the question every BGG tool gets asked
Eric has seen the criticism land on other BGG apps: why does this
thing want my password? The README now answers it where the
requirement appears: BGG has no write API, so uploading means signing
into the real website in a visible browser on the user's own machine
— that login is the password's entire job. And the reassurance that
matters: no server, no telemetry, no analytics, nothing collected;
credentials go to boardgamegeek.com and nowhere else, the only other
contact is the user's own chosen vision provider (photos only, and a
local Ollama keeps even those home). The guide's credentials section
links back and notes the saved browser session stays local too.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jXZFSTZQKzAC8fqpWSz9g
2026-08-06 10:06:43 -04:00

10 KiB

The bggpipe user's guide

Everything past the README's quick start: credentials, every stage and its flags, the web app, phones, RPGs, fixing the model's mistakes, and running without a BGG token. Screenshots of everything described here: the tour.

Contents

Credentials and configuration

The init wizard is idempotent — re-run it anytime to check status or add keys you skipped. It prompts for the credentials below (hidden input, saved to a .env it creates with owner-only permissions) and offers the one-time Playwright Chromium download. Prefer doing it by hand? Copy .env.example beside your data, fill it in, and run playwright install chromium yourself.

Secrets live in environment variables only, never in config files, code, or logs, and .env is gitignored. Every bggpipe command loads .env from the working directory by itself — real environment variables always win over the file, so direnv users and CI overrides keep working unchanged.

Your credentials never leave your machine except to sign in to boardgamegeek.com itself — the README spells out the full privacy picture. The password exists solely because BGG has no write API: adding games means driving the real website, in a visible browser window, on your computer. The saved browser session (storage_state.json) is credential-adjacent — it stays local and gitignored too.

Variable Used by What it is
ANTHROPIC_API_KEY extract Anthropic API key
BGG_API_TOKEN resolve, diff, enrich Bearer token from your registered BGG application
BGG_USERNAME diff, upload, enrich Your BGG username (public, but kept in .env so it lives in one place)
BGG_PASSWORD upload (website login) Your BGG password

Non-secret knobs live in config.toml: photos_dir, data_dir, the BGG rate limit, and the vision setup — a [vision.<provider>] block per provider ("anthropic" or any OpenAI-compatible endpoint, including a local Ollama), with vision_provider picking one. Local models read spines noticeably worse than frontier ones — expect a longer proofread pass on the Titles page, not a broken pipeline.

The web app

bggpipe web    # opens http://127.0.0.1:8377/ — the whole app in the browser

Seven pages — Pipeline, Photos, Titles, Review, Queue, Library, and Help — all pictured in the tour. Stage runs execute one at a time in the background with live output; every decision saves immediately; the pages live-follow the data files, so a stage run in another terminal shows up without a refresh. The real upload sits behind a confirmation (and behind a stub-data lock if synthetic test fixtures ever regenerate). The in-app Help page documents every status chip and keyboard shortcut.

From your phone

The app is localhost-only by default. To use it from a phone or tablet on your network — proofreading from the couch, or shooting shelf photos straight into the pipeline — serve it to the LAN instead:

bggpipe web --lan   # localhost + your network, behind an access key

Startup prints a pairing link (?k=...) and a QR code: point the phone's camera at the terminal and tap. Pairing is one-time per device — the key persists across restarts (data/.lan_key; delete it to revoke every paired device) and the cookie lasts a year. Save the page to the phone's home screen for the full-screen treatment, piper icon included.

To photograph shelves from the phone: on the Photos page, tap the drop zone and choose "Take Photo." The upload narrates its progress, and camera captures get unique shelf-<timestamp> names so rapid-fire shots never overwrite each other. The key is the only lock — there is no login behind it — so still prefer networks you trust (or use a device VPN like Tailscale against the localhost default instead).

The stages, from the terminal

Every stage is also a command, and the two interfaces share all state:

bggpipe extract              # photos → titles.json (+ retake prompts)
bggpipe resolve              # titles → BGG ids/versions in matches.csv
bggpipe review --web         # review UI only
bggpipe diff                 # compare against your BGG collection
bggpipe upload --dry-run     # ALWAYS inspect this first
bggpipe upload --limit 1     # then one game, then small batches
bggpipe enrich               # full metadata → data/games.json

Each stage skips work it has already done; --force/--refresh flags redo it. Every stage is idempotent and resumable: kill it mid-run, restart, lose nothing. All artifacts are flat CSV/JSON files you can inspect and edit. review without --web runs in the terminal instead.

Taking good shelf photos

Straight-on, one shelf (or part of one) per shot, close enough that spine text is legible to a human. If you can't read it, the model can't either. Overlap between shots is fine: duplicate reads are deduped automatically, with the merge shown (and veto-able) in review. Boxes the model spots but can't identify become retake prompts in unidentified.json and the Photos page's "reshoot" tickets: photograph those boxes up close, drop the new photo in, and run extract again.

Fixing what the model gets wrong

Vision reads aren't perfect, and you know things the photos don't show. The Titles page lets you edit a title (fix a misspelling, add publisher/edition/year/language cues you know offhand), split a line into per-photo copies when one title is actually several boxes, remove lines that aren't games at all, and add a game no photo caught. Every one of these is durable: the decision lands in a small committed store (data/title_edits.json, data/title_splits.json, data/title_removals.json, data/title_additions.json) that is replayed on every rebuild — re-running extract or resolve can never undo your curation. Undo any decision by deleting its record from the store.

Games BGG doesn't have at all can be kept as local library citizens: their detail page in the Library takes hand-written facts (players, playtime, publisher, notes) and a cover photo of your own, stored in data/local_games.json and data/local_art/ — the only source such a game will ever have.

RPGs on your shelves

Tabletop RPGs aren't in BGG's board-game database — they live on RPGGeek, which shares the same underlying API. When a title isn't found as a board game, bggpipe retries as an RPG: matches are identified, enriched (designers, publishers, genres from RPGGeek), and browsable in the library (filter: RPGs), but they stay local only — they're never uploaded, since your BGG collection can't hold them. When the automatic search can't reach the right database (BGG has board games named "Dungeons & Dragons" too), every Review card has explicit search BGG / search RPGGeek buttons.

Uploading safely

upload drives a real logged-in browser session against your real account, so it is deliberately careful:

  • --dry-run logs what would happen without touching the site — always read it first, then --limit 1, then small batches.
  • The browser runs headed by default — BGG's Cloudflare check blocks headless ones, and a first login may need one human click before the session is saved locally and reused.
  • Requests are slow on purpose (seconds between actions, per BGG's API policy); the Queue page shows exactly what will run before it runs, and upload_log.csv keeps a permanent record of every attempt.
  • --retry-failed re-attempts failures; --verify re-fetches your collection and cross-checks the log. Note that BGG's collection export can lag the website by hours — freshly-landed work may look missing to diff/--verify until it catches up.
  • Review decisions outrank the queue: re-deciding a match after diff retires its queued job automatically.

Running before your BGG token arrives

BGG application approval can take a week or more. Until then: extract works immediately (it only needs the vision key), and resolve does what it can, parking the rest as "waiting on BGG API token" — it picks them up automatically once the token exists. Everything is saved as you go.

diff normally fetches your collection live, but there's a logged-in-browser exemption that needs no token: while signed in to BGG, save these two URLs as data/collection_snapshot_base.xml and data/collection_snapshot_expansions.xml (if you get a "queued" message, refresh after a few seconds):

  • https://boardgamegeek.com/xmlapi2/collection?username=YOU&own=1&version=1
  • https://boardgamegeek.com/xmlapi2/collection?username=YOU&own=1&version=1&subtype=boardgameexpansion

Keeping your data safe from git

The README's quick start — installed as a tool, run in a directory of your own — is the only supported way to use bggpipe on your collection. Everything the pipeline produces lives where you run it, and uv tool upgrade bggpipe picks up fixes without going anywhere near your data.

Why not clone and run? The source repo doubles as its author's live pipeline: data/ ships with their real artifacts, committed and updated often. Run inside a clone and your data lands at git-tracked paths — the next git pull will refuse to merge, and the usual remedies (git reset --hard, git checkout ., git stash, git clean -fdx) would destroy your review decisions, hand-written games, upload log, and photos. bggpipe detects this arrangement and warns at init and on the web dashboard; don't ignore it.

One committed file of the author's data deserves a word: data/STUB_DATA.marker is normally absent. It appears only if the synthetic stub fixtures (from scripts/write_stub_fixtures.py) regenerate the CSVs, and the upload stage refuses to run while it exists — so placeholder data can never reach a real BGG account.