Extract stage: vision title + edition-cue extraction, offline-tested

Per-photo raw results cached under data/extract_raw/ (gitignored) so
re-runs are free, --only re-extracts a single photo, and titles.json is
rebuilt with dedupe that keeps conflicting-edition sightings separate.
HEIC converts via pillow-heif; images downscale to <=1568px long edge;
model JSON parsed defensively (code fences, surrounding prose). Vision
callable is injectable — tests use a local fake; the real one uses the
anthropic SDK (approved) with the model from config.toml.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Eric Wagoner
2026-08-01 13:10:13 -04:00
parent 96a1e09956
commit b7a1ef8549
6 changed files with 803 additions and 1 deletions
+1
View File
@@ -7,6 +7,7 @@ storage_state.json
# Local inputs & cache (CSV/JSON artifacts in data/ ARE committed)
photos/
data/bgg_cache/
data/extract_raw/
# Python
__pycache__/