The README stops being four documents wearing one trench coat
Eric's read on the first-visitor experience: 190 lines of pitch, manual, gallery, and contributor doc is intimidating when the visitor only needs the first 40. Split three ways: README.md is now the front door — what it is, why it exists (told in first person now, since it IS a personal itch scratched), how the six stages work, requirements, quick start, one hero screenshot, and the development/citizenship/license notes. Sixty percent shorter. docs/tour.md carries the full gallery: all seven pages, the game detail view, and the phone set, captions intact. docs/guide.md is the complete user's guide: credentials and config, the stages and their flags, phone pairing, photo technique, curation stores, RPG handling, upload safety (including the collection-export lag), the no-token-yet path, and the keep-data-out-of-git rationale. Every relative link and README→guide anchor machine-verified to resolve. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jXZFSTZQKzAC8fqpWSz9g
This commit is contained in:
co-authored by
Claude Fable 5
parent
1df784e253
commit
6041150c8d
+108
@@ -0,0 +1,108 @@
|
||||
# The bggpipe user's guide
|
||||
|
||||
Everything past the [README](../README.md)'s quick start: credentials, every stage and its flags, the web app, phones, RPGs, fixing the model's mistakes, and running without a BGG token. Screenshots of everything described here: the [tour](tour.md).
|
||||
|
||||
## Contents
|
||||
|
||||
- [Credentials and configuration](#credentials-and-configuration)
|
||||
- [The web app](#the-web-app)
|
||||
- [From your phone](#from-your-phone)
|
||||
- [The stages, from the terminal](#the-stages-from-the-terminal)
|
||||
- [Taking good shelf photos](#taking-good-shelf-photos)
|
||||
- [Fixing what the model gets wrong](#fixing-what-the-model-gets-wrong)
|
||||
- [RPGs on your shelves](#rpgs-on-your-shelves)
|
||||
- [Uploading safely](#uploading-safely)
|
||||
- [Running before your BGG token arrives](#running-before-your-bgg-token-arrives)
|
||||
- [Keeping your data safe from git](#keeping-your-data-safe-from-git)
|
||||
|
||||
## Credentials and configuration
|
||||
|
||||
The `init` wizard is idempotent — re-run it anytime to check status or add keys you skipped. It prompts for the credentials below (hidden input, saved to a `.env` it creates with owner-only permissions) and offers the one-time Playwright Chromium download. Prefer doing it by hand? Copy [.env.example](../.env.example) beside your data, fill it in, and run `playwright install chromium` yourself.
|
||||
|
||||
Secrets live in environment variables only, never in config files, code, or logs, and `.env` is gitignored. Every `bggpipe` command loads `.env` from the working directory by itself — real environment variables always win over the file, so [direnv](https://direnv.net/) users and CI overrides keep working unchanged.
|
||||
|
||||
| Variable | Used by | What it is |
|
||||
|---|---|---|
|
||||
| `ANTHROPIC_API_KEY` | extract | Anthropic API key |
|
||||
| `BGG_API_TOKEN` | resolve, diff, enrich | Bearer token from your registered BGG application |
|
||||
| `BGG_USERNAME` | diff, upload, enrich | Your BGG username (public, but kept in `.env` so it lives in one place) |
|
||||
| `BGG_PASSWORD` | upload (website login) | Your BGG password |
|
||||
|
||||
Non-secret knobs live in `config.toml`: `photos_dir`, `data_dir`, the BGG rate limit, and the vision setup — a `[vision.<provider>]` block per provider ("anthropic" or any OpenAI-compatible endpoint, including a local [Ollama](https://ollama.com/)), with `vision_provider` picking one. Local models read spines noticeably worse than frontier ones — expect a longer proofread pass on the Titles page, not a broken pipeline.
|
||||
|
||||
## The web app
|
||||
|
||||
```sh
|
||||
bggpipe web # opens http://127.0.0.1:8377/ — the whole app in the browser
|
||||
```
|
||||
|
||||
Seven pages — Pipeline, Photos, Titles, Review, Queue, Library, and Help — all [pictured in the tour](tour.md). Stage runs execute one at a time in the background with live output; every decision saves immediately; the pages live-follow the data files, so a stage run in another terminal shows up without a refresh. The real upload sits behind a confirmation (and behind a stub-data lock if synthetic test fixtures ever regenerate). The in-app **Help** page documents every status chip and keyboard shortcut.
|
||||
|
||||
## From your phone
|
||||
|
||||
The app is localhost-only by default. To use it from a phone or tablet on your network — proofreading from the couch, or shooting shelf photos straight into the pipeline — serve it to the LAN instead:
|
||||
|
||||
```sh
|
||||
bggpipe web --lan # localhost + your network, behind an access key
|
||||
```
|
||||
|
||||
Startup prints a pairing link (`?k=...`) and a QR code: point the phone's camera at the terminal and tap. Pairing is one-time per device — the key persists across restarts (`data/.lan_key`; delete it to revoke every paired device) and the cookie lasts a year. Save the page to the phone's home screen for the full-screen treatment, piper icon included.
|
||||
|
||||
To photograph shelves from the phone: on the Photos page, tap the drop zone and choose "Take Photo." The upload narrates its progress, and camera captures get unique `shelf-<timestamp>` names so rapid-fire shots never overwrite each other. The key is the only lock — there is no login behind it — so still prefer networks you trust (or use a device VPN like Tailscale against the localhost default instead).
|
||||
|
||||
## The stages, from the terminal
|
||||
|
||||
Every stage is also a command, and the two interfaces share all state:
|
||||
|
||||
```sh
|
||||
bggpipe extract # photos → titles.json (+ retake prompts)
|
||||
bggpipe resolve # titles → BGG ids/versions in matches.csv
|
||||
bggpipe review --web # review UI only
|
||||
bggpipe diff # compare against your BGG collection
|
||||
bggpipe upload --dry-run # ALWAYS inspect this first
|
||||
bggpipe upload --limit 1 # then one game, then small batches
|
||||
bggpipe enrich # full metadata → data/games.json
|
||||
```
|
||||
|
||||
Each stage skips work it has already done; `--force`/`--refresh` flags redo it. Every stage is idempotent and resumable: kill it mid-run, restart, lose nothing. All artifacts are flat CSV/JSON files you can inspect and edit. `review` without `--web` runs in the terminal instead.
|
||||
|
||||
## Taking good shelf photos
|
||||
|
||||
Straight-on, one shelf (or part of one) per shot, close enough that spine text is legible to a human. If you can't read it, the model can't either. Overlap between shots is fine: duplicate reads are deduped automatically, with the merge shown (and veto-able) in review. Boxes the model spots but can't identify become retake prompts in `unidentified.json` and the Photos page's "reshoot" tickets: photograph those boxes up close, drop the new photo in, and run `extract` again.
|
||||
|
||||
## Fixing what the model gets wrong
|
||||
|
||||
Vision reads aren't perfect, and you know things the photos don't show. The Titles page lets you **edit** a title (fix a misspelling, add publisher/edition/year/language cues you know offhand), **split** a line into per-photo copies when one title is actually several boxes, **remove** lines that aren't games at all, and **add** a game no photo caught. Every one of these is durable: the decision lands in a small committed store (`data/title_edits.json`, `data/title_splits.json`, `data/title_removals.json`, `data/title_additions.json`) that is replayed on every rebuild — re-running extract or resolve can never undo your curation. Undo any decision by deleting its record from the store.
|
||||
|
||||
Games BGG doesn't have at all can be kept as **local** library citizens: their detail page in the Library takes hand-written facts (players, playtime, publisher, notes) and a cover photo of your own, stored in `data/local_games.json` and `data/local_art/` — the only source such a game will ever have.
|
||||
|
||||
## RPGs on your shelves
|
||||
|
||||
Tabletop RPGs aren't in BGG's board-game database — they live on RPGGeek, which shares the same underlying API. When a title isn't found as a board game, bggpipe retries as an RPG: matches are identified, enriched (designers, publishers, genres from RPGGeek), and browsable in the library (filter: RPGs), but they stay **local only** — they're never uploaded, since your BGG collection can't hold them. When the automatic search can't reach the right database (BGG has board games named "Dungeons & Dragons" too), every Review card has explicit **search BGG** / **search RPGGeek** buttons.
|
||||
|
||||
## Uploading safely
|
||||
|
||||
`upload` drives a real logged-in browser session against your real account, so it is deliberately careful:
|
||||
|
||||
- `--dry-run` logs what would happen without touching the site — always read it first, then `--limit 1`, then small batches.
|
||||
- The browser runs **headed** by default — BGG's Cloudflare check blocks headless ones, and a first login may need one human click before the session is saved locally and reused.
|
||||
- Requests are slow on purpose (seconds between actions, per BGG's API policy); the Queue page shows exactly what will run before it runs, and `upload_log.csv` keeps a permanent record of every attempt.
|
||||
- `--retry-failed` re-attempts failures; `--verify` re-fetches your collection and cross-checks the log. Note that BGG's collection export can lag the website by hours — freshly-landed work may look missing to `diff`/`--verify` until it catches up.
|
||||
- Review decisions outrank the queue: re-deciding a match after `diff` retires its queued job automatically.
|
||||
|
||||
## Running before your BGG token arrives
|
||||
|
||||
BGG application approval can take a week or more. Until then: `extract` works immediately (it only needs the vision key), and `resolve` does what it can, parking the rest as "waiting on BGG API token" — it picks them up automatically once the token exists. Everything is saved as you go.
|
||||
|
||||
`diff` normally fetches your collection live, but there's a logged-in-browser exemption that needs no token: while signed in to BGG, save these two URLs as `data/collection_snapshot_base.xml` and `data/collection_snapshot_expansions.xml` (if you get a "queued" message, refresh after a few seconds):
|
||||
|
||||
- `https://boardgamegeek.com/xmlapi2/collection?username=YOU&own=1&version=1`
|
||||
- `https://boardgamegeek.com/xmlapi2/collection?username=YOU&own=1&version=1&subtype=boardgameexpansion`
|
||||
|
||||
## Keeping your data safe from git
|
||||
|
||||
The [README's quick start](../README.md#quick-start) — installed as a tool, run in a directory of your own — is the only supported way to use bggpipe on your collection. Everything the pipeline produces lives where you run it, and `uv tool upgrade bggpipe` picks up fixes without going anywhere near your data.
|
||||
|
||||
**Why not clone and run?** The source repo doubles as its author's live pipeline: `data/` ships with their real artifacts, committed and updated often. Run inside a clone and *your* data lands at git-tracked paths — the next `git pull` will refuse to merge, and the usual remedies (`git reset --hard`, `git checkout .`, `git stash`, `git clean -fdx`) would destroy your review decisions, hand-written games, upload log, and photos. bggpipe detects this arrangement and warns at `init` and on the web dashboard; don't ignore it.
|
||||
|
||||
One committed file of the author's data deserves a word: `data/STUB_DATA.marker` is normally absent. It appears only if the synthetic stub fixtures (from `scripts/write_stub_fixtures.py`) regenerate the CSVs, and the upload stage refuses to run while it exists — so placeholder data can never reach a real BGG account.
|
||||
Reference in New Issue
Block a user