Extract stage: vision title + edition-cue extraction, offline-tested

Per-photo raw results cached under data/extract_raw/ (gitignored) so
re-runs are free, --only re-extracts a single photo, and titles.json is
rebuilt with dedupe that keeps conflicting-edition sightings separate.
HEIC converts via pillow-heif; images downscale to <=1568px long edge;
model JSON parsed defensively (code fences, surrounding prose). Vision
callable is injectable — tests use a local fake; the real one uses the
anthropic SDK (approved) with the model from config.toml.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Eric Wagoner
2026-08-01 13:10:13 -04:00
parent 96a1e09956
commit b7a1ef8549
6 changed files with 803 additions and 1 deletions
+4
View File
@@ -8,6 +8,10 @@ dependencies = [
"httpx>=0.27",
"rapidfuzz>=3.9",
"defusedxml>=0.7.1",
"anthropic>=0.120.2",
"pillow>=12.3.0",
"pillow-heif>=1.5.0",
"rich>=15.0.0",
]
[project.scripts]