Audit round 5 (curation feature): 5 blind reviewers, 14 confirmed fixes
The standing post-feature audit over a7f0cfe. Correctness (data): splits
become photo-scoped store records so splitting one edition no longer
force-splits same-named editions, and renaming a split copy migrates its
protection to the corrected title instead of silently re-merging copies.
Correctness (web): edit scoping now counts siblings by NORMALIZED title
(matching how stored edits apply), same-title-same-photos edits are
refused rather than corrupting the sibling entry, split copies serve
their real per-photo cues to the edit form instead of blanks, and a
split whose row vanished underneath returns 409 instead of a false 200.
Silent failures: replay_titles refuses to rebuild from a PARTIAL raw
cache (fresh clone + one --only extract would have truncated the
committed titles.json); the edit endpoint writes in crash-safe order
(cull, record, replay); corrupt curation stores fail loud naming the
file; retried edits don't double-record. Review-decision durability:
drop_rows never drops dedupe_veto rows — a rename retitles them in
place — and writes through a no-reload path so a concurrent rewrite
can't silently discard the cull. Style: catalog action cells get their
own class (.rowactions' flex display broke table alignment), editor
inputs match the design system and stop overriding the global
focus-visible outline, EditBody's clear-semantics docstring scoped to
cue fields, "nothing to change" derived from the record itself.
Tests: 8 new (photo-scoped splits, veto preservation, photo-narrowed
drops, 409s on both curation endpoints under a running job, partial-raw
replay guard, rename-keeps-protection lifecycle, corrupt-store error,
cue-field editing) and the dead edition_hint key in the edit test now
exercises real cue fields. 259 passing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jXZFSTZQKzAC8fqpWSz9g
This commit is contained in:
@@ -47,7 +47,7 @@ Full design lives in `bgg-shelf-pipeline-spec.md` (read it before changing pipel
|
|||||||
- Base game vs. expansion vs. new edition is the top failure mode — bias matching toward `ambiguous` over auto-match ("Wingspan Europe" must not match base Wingspan).
|
- Base game vs. expansion vs. new edition is the top failure mode — bias matching toward `ambiguous` over auto-match ("Wingspan Europe" must not match base Wingspan).
|
||||||
- Editions/versions matter: Eric owns multiple editions of some games — each is a separate collection entry (keyed by `collid` on BGG). Never guess a version: no legible cues → `version_unknown` and a version-less collection entry.
|
- Editions/versions matter: Eric owns multiple editions of some games — each is a separate collection entry (keyed by `collid` on BGG). Never guess a version: no legible cues → `version_unknown` and a version-less collection entry.
|
||||||
- Normalize titles (casefold, strip punctuation/articles, special chars like é/&/:) identically on both sides of a match; dedupe across photos but keep `source_photos` provenance.
|
- Normalize titles (casefold, strip punctuation/articles, special chars like é/&/:) identically on both sides of a match; dedupe across photos but keep `source_photos` provenance.
|
||||||
- **Human curation is durable**: `data/title_splits.json` (titles split into per-photo copies) and `data/title_edits.json` (corrected reads/cues) are replayed on every titles.json rebuild, and both extract's and resolve's dedupe honor them forever. Row-level split decisions also persist via the `dedupe_veto` column.
|
- **Human curation is durable**: `data/title_splits.json` (photo-scoped split-into-copies decisions, honored by extract's dedupe AND resolve's dedupe) and `data/title_edits.json` (corrected reads/cues, applied before dedupe on every titles.json rebuild) persist forever. Row-level decisions persist via the `dedupe_veto` column — edits never drop veto'd rows (a rename retitles them in place).
|
||||||
- **RPGs are local-only citizens**: when the board-game search runs dry, resolve falls back to `type=rpgitem` (same geekdo API/token). Matched rpgitems enrich into the library but diff routes them to `local_only` — they must never reach `to_add.csv`/upload (their collection lives on RPGGeek, out of scope).
|
- **RPGs are local-only citizens**: when the board-game search runs dry, resolve falls back to `type=rpgitem` (same geekdo API/token). Matched rpgitems enrich into the library but diff routes them to `local_only` — they must never reach `to_add.csv`/upload (their collection lives on RPGGeek, out of scope).
|
||||||
- Detailed BGG API behavior (202 queueing, collection-endpoint quirks, endpoints): use the `bgg-api` skill. **If the spec's BGG behavior changes, update the `bgg-api` skill to match** — they must not drift.
|
- Detailed BGG API behavior (202 queueing, collection-endpoint quirks, endpoints): use the `bgg-api` skill. **If the spec's BGG behavior changes, update the `bgg-api` skill to match** — they must not drift.
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,3 @@
|
|||||||
|
[
|
||||||
|
"Wiz-War"
|
||||||
|
]
|
||||||
+30
-6
@@ -260,16 +260,14 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"title_raw": "Wiz-War",
|
"title_raw": "Wiz-War",
|
||||||
"confidence": "high",
|
"confidence": "medium",
|
||||||
"publisher_hint": "Fantasy Flight Games",
|
"publisher_hint": "",
|
||||||
"edition_hint": "9th Edition",
|
"edition_hint": "",
|
||||||
"year_hint": null,
|
"year_hint": null,
|
||||||
"language_hint": "English",
|
"language_hint": "English",
|
||||||
"art_notes": "black spine with purple/blue text, partial visible 'Wiz-Wa...' with subtitle 'and board game of ...ly battle and treasures!'",
|
"art_notes": "black spine with purple/blue text, partial visible 'Wiz-Wa...' with subtitle 'and board game of ...ly battle and treasures!'",
|
||||||
"source_photos": [
|
"source_photos": [
|
||||||
"IMG_4502.jpeg",
|
"IMG_4502.jpeg"
|
||||||
"IMG_4504.jpeg",
|
|
||||||
"IMG_4528.jpeg"
|
|
||||||
],
|
],
|
||||||
"title_normalized": "wiz war"
|
"title_normalized": "wiz war"
|
||||||
},
|
},
|
||||||
@@ -351,6 +349,19 @@
|
|||||||
],
|
],
|
||||||
"title_normalized": "dungeon"
|
"title_normalized": "dungeon"
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"title_raw": "WIZ-WAR",
|
||||||
|
"confidence": "high",
|
||||||
|
"publisher_hint": "Fantasy Flight Games",
|
||||||
|
"edition_hint": "",
|
||||||
|
"year_hint": null,
|
||||||
|
"language_hint": "English",
|
||||||
|
"art_notes": "Dark spine with wizard/warrior artwork, sepia tones",
|
||||||
|
"source_photos": [
|
||||||
|
"IMG_4504.jpeg"
|
||||||
|
],
|
||||||
|
"title_normalized": "wiz war"
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"title_raw": "TICKET TO RIDE",
|
"title_raw": "TICKET TO RIDE",
|
||||||
"confidence": "high",
|
"confidence": "high",
|
||||||
@@ -1280,6 +1291,19 @@
|
|||||||
],
|
],
|
||||||
"title_normalized": "red dragon inn smorgasbox"
|
"title_normalized": "red dragon inn smorgasbox"
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"title_raw": "WIZ-WAR",
|
||||||
|
"confidence": "high",
|
||||||
|
"publisher_hint": "",
|
||||||
|
"edition_hint": "9th Edition",
|
||||||
|
"year_hint": null,
|
||||||
|
"language_hint": "English",
|
||||||
|
"art_notes": "Dark purple/maroon box with pink outlined logo text, illustration of a turbaned wizard character casting fire spell in bottom right corner, tagline 'KILL THEM WITH FIRE!'",
|
||||||
|
"source_photos": [
|
||||||
|
"IMG_4528.jpeg"
|
||||||
|
],
|
||||||
|
"title_normalized": "wiz war"
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"title_raw": "SLUGFEST GAMES",
|
"title_raw": "SLUGFEST GAMES",
|
||||||
"confidence": "low",
|
"confidence": "low",
|
||||||
|
|||||||
+84
-27
@@ -251,18 +251,33 @@ EDIT_FIELDS = (
|
|||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _load_store(path: Path) -> list:
|
||||||
|
"""Curation stores hold irreplaceable human decisions and are committed
|
||||||
|
(merge conflicts are a realistic corruption vector) — so a broken file
|
||||||
|
must stop the pipeline loudly, never quietly reset it."""
|
||||||
|
if not path.exists():
|
||||||
|
return []
|
||||||
|
try:
|
||||||
|
return json.loads(path.read_text())
|
||||||
|
except json.JSONDecodeError as err:
|
||||||
|
raise ValueError(
|
||||||
|
f"{path} is corrupt ({err}) — fix or delete it; it holds human "
|
||||||
|
"review decisions, so check git history before deleting"
|
||||||
|
) from err
|
||||||
|
|
||||||
|
|
||||||
def load_title_edits(path: Path) -> list[dict]:
|
def load_title_edits(path: Path) -> list[dict]:
|
||||||
"""Human corrections to raw reads (fixed misspellings, cues the owner
|
"""Human corrections to raw reads (fixed misspellings, cues the owner
|
||||||
knows offhand). Each record: {"match": <title as displayed when the fix
|
knows offhand). Each record: {"match": <title as displayed when the fix
|
||||||
was made>, "photos": [...] to target one copy (optional), <EDIT_FIELDS
|
was made>, "photos": [...] to target one copy (optional), <EDIT_FIELDS
|
||||||
to override>}. Applied on every rebuild, before dedupe."""
|
to override>}. Applied on every rebuild, before dedupe."""
|
||||||
if not path.exists():
|
return _load_store(path)
|
||||||
return []
|
|
||||||
return json.loads(path.read_text())
|
|
||||||
|
|
||||||
|
|
||||||
def record_title_edit(path: Path, record: dict) -> None:
|
def record_title_edit(path: Path, record: dict) -> None:
|
||||||
existing = load_title_edits(path)
|
existing = load_title_edits(path)
|
||||||
|
if existing and existing[-1] == record:
|
||||||
|
return # a retried request must not double-record
|
||||||
atomic_write_text(
|
atomic_write_text(
|
||||||
path, json.dumps([*existing, record], indent=2, ensure_ascii=False) + "\n"
|
path, json.dumps([*existing, record], indent=2, ensure_ascii=False) + "\n"
|
||||||
)
|
)
|
||||||
@@ -298,39 +313,69 @@ def apply_title_edits(entries: list[dict], edits: list[dict]) -> list[dict]:
|
|||||||
return entries
|
return entries
|
||||||
|
|
||||||
|
|
||||||
def load_title_splits(path: Path) -> set[str]:
|
def load_title_splits(path: Path) -> list[dict]:
|
||||||
"""Normalized titles the human split into per-photo copies — a durable
|
"""The human's split-into-copies decisions — durable: they must survive
|
||||||
review decision that must survive extract rebuilds and resolve --force."""
|
extract rebuilds and resolve --force. Each stored record is
|
||||||
if not path.exists():
|
{"title": ..., "photos": [...]} scoping the split to the sightings that
|
||||||
return set()
|
were on the split line (a bare string is legacy: unscoped). Returns
|
||||||
return {normalize_title(t) for t in json.loads(path.read_text())}
|
[{"norm": <normalized title>, "photos": set | None}]."""
|
||||||
|
records = []
|
||||||
|
for item in _load_store(path):
|
||||||
|
if isinstance(item, str):
|
||||||
|
records.append({"norm": normalize_title(item), "photos": None})
|
||||||
|
else:
|
||||||
|
photos = item.get("photos")
|
||||||
|
records.append(
|
||||||
|
{
|
||||||
|
"norm": normalize_title(item["title"]),
|
||||||
|
"photos": set(photos) if photos else None,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return records
|
||||||
|
|
||||||
|
|
||||||
def record_title_split(path: Path, title: str) -> None:
|
def is_split(norm: str, photos, splits: list[dict]) -> bool:
|
||||||
existing = json.loads(path.read_text()) if path.exists() else []
|
"""Does a split decision cover this (normalized title, photo set)?
|
||||||
if normalize_title(title) in {normalize_title(t) for t in existing}:
|
Photo-scoped records only bind sightings from the photos that were on
|
||||||
return
|
the split line — a same-named different edition keeps deduping."""
|
||||||
atomic_write_text(
|
photo_set = set(photos or [])
|
||||||
path, json.dumps([*existing, title], indent=2, ensure_ascii=False) + "\n"
|
return any(
|
||||||
|
r["norm"] == norm and (r["photos"] is None or r["photos"] & photo_set)
|
||||||
|
for r in splits
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
def dedupe_entries(
|
def record_title_split(path: Path, title: str, photos: list[str] | None = None) -> None:
|
||||||
entries: list[dict], split_titles: set[str] = frozenset()
|
if is_split(normalize_title(title), photos or [], load_title_splits(path)):
|
||||||
) -> list[dict]:
|
return
|
||||||
|
existing = _load_store(path)
|
||||||
|
record = {"title": title, "photos": sorted(photos) if photos else None}
|
||||||
|
atomic_write_text(
|
||||||
|
path, json.dumps([*existing, record], indent=2, ensure_ascii=False) + "\n"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def dedupe_entries(entries: list[dict], splits: list[dict] | None = None) -> list[dict]:
|
||||||
"""Collapse same-normalized-title sightings unless their cues conflict —
|
"""Collapse same-normalized-title sightings unless their cues conflict —
|
||||||
or unless the human declared the title split (several physical copies):
|
or unless the human declared them split (several physical copies):
|
||||||
those stay one entry per photo."""
|
split sightings stay one entry per photo and never absorb new ones."""
|
||||||
|
splits = splits or []
|
||||||
|
|
||||||
|
def covered(e: dict) -> bool:
|
||||||
|
return is_split(e["title_normalized"], e.get("source_photos"), splits)
|
||||||
|
|
||||||
result: list[dict] = []
|
result: list[dict] = []
|
||||||
for entry in entries:
|
for entry in entries:
|
||||||
entry = {**entry, "title_normalized": normalize_title(entry["title_raw"])}
|
entry = {**entry, "title_normalized": normalize_title(entry["title_raw"])}
|
||||||
if entry["title_normalized"] in split_titles:
|
if covered(entry):
|
||||||
result.append(entry)
|
result.append(entry)
|
||||||
continue
|
continue
|
||||||
for existing in result:
|
for existing in result:
|
||||||
if existing["title_normalized"] == entry[
|
if (
|
||||||
"title_normalized"
|
existing["title_normalized"] == entry["title_normalized"]
|
||||||
] and not cues_conflict(existing, entry):
|
and not covered(existing)
|
||||||
|
and not cues_conflict(existing, entry)
|
||||||
|
):
|
||||||
existing.update(_merge(existing, entry))
|
existing.update(_merge(existing, entry))
|
||||||
break
|
break
|
||||||
else:
|
else:
|
||||||
@@ -342,7 +387,7 @@ def rebuild_artifacts(
|
|||||||
raw_dir: Path,
|
raw_dir: Path,
|
||||||
titles_path: Path,
|
titles_path: Path,
|
||||||
unidentified_path: Path,
|
unidentified_path: Path,
|
||||||
split_titles: set[str] = frozenset(),
|
splits: list[dict] | None = None,
|
||||||
edits: list[dict] | None = None,
|
edits: list[dict] | None = None,
|
||||||
) -> tuple[list[dict], dict[str, list[dict]]]:
|
) -> tuple[list[dict], dict[str, list[dict]]]:
|
||||||
"""Regenerate titles.json and unidentified.json from the per-photo raw
|
"""Regenerate titles.json and unidentified.json from the per-photo raw
|
||||||
@@ -365,7 +410,7 @@ def rebuild_artifacts(
|
|||||||
photo = raw_file.name.removesuffix(".json")
|
photo = raw_file.name.removesuffix(".json")
|
||||||
if data.get("unidentified"):
|
if data.get("unidentified"):
|
||||||
unidentified[photo] = data["unidentified"]
|
unidentified[photo] = data["unidentified"]
|
||||||
deduped = dedupe_entries(apply_title_edits(entries, edits or []), split_titles)
|
deduped = dedupe_entries(apply_title_edits(entries, edits or []), splits)
|
||||||
titles_path.parent.mkdir(parents=True, exist_ok=True)
|
titles_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
atomic_write_text(
|
atomic_write_text(
|
||||||
titles_path, json.dumps(deduped, indent=2, ensure_ascii=False) + "\n"
|
titles_path, json.dumps(deduped, indent=2, ensure_ascii=False) + "\n"
|
||||||
@@ -385,7 +430,19 @@ def replay_titles(cfg: Config) -> None:
|
|||||||
splits = load_title_splits(cfg.title_splits_path)
|
splits = load_title_splits(cfg.title_splits_path)
|
||||||
edits = load_title_edits(cfg.title_edits_path)
|
edits = load_title_edits(cfg.title_edits_path)
|
||||||
raw_dir = cfg.extract_raw_dir
|
raw_dir = cfg.extract_raw_dir
|
||||||
if raw_dir.is_dir() and any(raw_dir.glob("*.json")):
|
raw_photos = (
|
||||||
|
{f.name.removesuffix(".json") for f in raw_dir.glob("*.json")}
|
||||||
|
if raw_dir.is_dir()
|
||||||
|
else set()
|
||||||
|
)
|
||||||
|
known_photos: set[str] = set()
|
||||||
|
if cfg.titles_path.exists():
|
||||||
|
for entry in json.loads(cfg.titles_path.read_text()):
|
||||||
|
known_photos.update(entry.get("source_photos") or [])
|
||||||
|
# raw caches are gitignored: a fresh clone (or a killed/partial extract)
|
||||||
|
# can have FEWER raw files than titles.json has photos — rebuilding from
|
||||||
|
# them would silently truncate the committed catalog
|
||||||
|
if raw_photos and raw_photos >= known_photos:
|
||||||
rebuild_artifacts(
|
rebuild_artifacts(
|
||||||
raw_dir, cfg.titles_path, cfg.unidentified_path, splits, edits
|
raw_dir, cfg.titles_path, cfg.unidentified_path, splits, edits
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -22,7 +22,7 @@ from rapidfuzz import fuzz
|
|||||||
|
|
||||||
from bggpipe.bgg_client import BGGAuthError, BGGClient, BGGQueueTimeout, client_for
|
from bggpipe.bgg_client import BGGAuthError, BGGClient, BGGQueueTimeout, client_for
|
||||||
from bggpipe.config import Config
|
from bggpipe.config import Config
|
||||||
from bggpipe.extract import cues_conflict, load_title_splits
|
from bggpipe.extract import cues_conflict, is_split, load_title_splits
|
||||||
from bggpipe.fsio import atomic_write_csv
|
from bggpipe.fsio import atomic_write_csv
|
||||||
from bggpipe.models import (
|
from bggpipe.models import (
|
||||||
RECOGNIZED_MATCH_STATUSES,
|
RECOGNIZED_MATCH_STATUSES,
|
||||||
@@ -431,7 +431,7 @@ class MergeEvent:
|
|||||||
def dedupe_matches(
|
def dedupe_matches(
|
||||||
rows: list[dict],
|
rows: list[dict],
|
||||||
titles: list[TitleEntry],
|
titles: list[TitleEntry],
|
||||||
split_titles: set[str] = frozenset(),
|
splits: list[dict] | None = None,
|
||||||
) -> list[MergeEvent]:
|
) -> list[MergeEvent]:
|
||||||
"""Post-resolve dedupe: rows resolving to the same (bgg_id, version_id —
|
"""Post-resolve dedupe: rows resolving to the same (bgg_id, version_id —
|
||||||
or both version-unknown) are the same physical game seen twice (a typo
|
or both version-unknown) are the same physical game seen twice (a typo
|
||||||
@@ -466,8 +466,12 @@ def dedupe_matches(
|
|||||||
# re-running resolve must never overturn that (spec: re-runs
|
# re-running resolve must never overturn that (spec: re-runs
|
||||||
# lose no work, least of all review decisions)
|
# lose no work, least of all review decisions)
|
||||||
continue
|
continue
|
||||||
if normalize_title(row["title_raw"]) in split_titles:
|
if is_split(
|
||||||
continue # human-split title: per-photo rows stay separate
|
normalize_title(row["title_raw"]),
|
||||||
|
row["source_photos"].split(";"),
|
||||||
|
splits or [],
|
||||||
|
):
|
||||||
|
continue # human-split copies: per-photo rows stay separate
|
||||||
key = (
|
key = (
|
||||||
row["bgg_id"],
|
row["bgg_id"],
|
||||||
row["version_id"] if is_confident_version(row) else "",
|
row["version_id"] if is_confident_version(row) else "",
|
||||||
|
|||||||
+36
-15
@@ -28,6 +28,7 @@ from bggpipe.models import (
|
|||||||
UNDECIDED_MATCH_STATUSES,
|
UNDECIDED_MATCH_STATUSES,
|
||||||
BGGResponseError,
|
BGGResponseError,
|
||||||
)
|
)
|
||||||
|
from bggpipe.normalize import normalize_title
|
||||||
from bggpipe.resolve import (
|
from bggpipe.resolve import (
|
||||||
MatchRow,
|
MatchRow,
|
||||||
TitleEntry,
|
TitleEntry,
|
||||||
@@ -158,6 +159,12 @@ class ReviewSession:
|
|||||||
"you decided — decision NOT saved"
|
"you decided — decision NOT saved"
|
||||||
)
|
)
|
||||||
return
|
return
|
||||||
|
self._write_rows()
|
||||||
|
|
||||||
|
def _write_rows(self) -> None:
|
||||||
|
"""Atomic write of self.rows exactly as they stand — no reload, so
|
||||||
|
a caller that just derived self.rows from a fresh reload can't have
|
||||||
|
its result silently replaced before the write."""
|
||||||
try:
|
try:
|
||||||
own_mtime = write_matches(self.cfg.matches_path, self.rows)
|
own_mtime = write_matches(self.cfg.matches_path, self.rows)
|
||||||
except OSError:
|
except OSError:
|
||||||
@@ -258,25 +265,39 @@ class ReviewSession:
|
|||||||
self._save(copies[0])
|
self._save(copies[0])
|
||||||
return copies
|
return copies
|
||||||
|
|
||||||
def drop_rows(self, title_raw: str, photos: list[str] | None = None) -> int:
|
def drop_rows(
|
||||||
|
self,
|
||||||
|
title_raw: str,
|
||||||
|
photos: list[str] | None = None,
|
||||||
|
new_title: str | None = None,
|
||||||
|
) -> int:
|
||||||
"""An edit invalidated these rows — the BGG match was made against
|
"""An edit invalidated these rows — the BGG match was made against
|
||||||
the uncorrected read. Remove them so resolve re-queries with the
|
the uncorrected read. Remove them so resolve re-queries with the
|
||||||
fix; `photos` narrows the cull to one copy of a split title."""
|
fix; `photos` narrows the cull to one copy of a split title.
|
||||||
|
Rows carrying a human veto (dedupe_veto) are never dropped — the
|
||||||
|
match itself was human-vetted, and dropping would erase the veto;
|
||||||
|
a rename updates their title in place so they follow the entry."""
|
||||||
self.reload_if_changed()
|
self.reload_if_changed()
|
||||||
|
norm = normalize_title(title_raw)
|
||||||
def stale(row: dict) -> bool:
|
keep: list[dict] = []
|
||||||
if row["title_raw"] != title_raw:
|
dropped = renamed = 0
|
||||||
return False
|
for row in self.rows:
|
||||||
if photos is None:
|
targeted = normalize_title(row["title_raw"]) == norm and (
|
||||||
return True
|
photos is None
|
||||||
row_photos = {p for p in row["source_photos"].split(";") if p}
|
or bool({p for p in row["source_photos"].split(";") if p} & set(photos))
|
||||||
return bool(row_photos & set(photos))
|
)
|
||||||
|
if not targeted:
|
||||||
keep = [r for r in self.rows if not stale(r)]
|
keep.append(row)
|
||||||
dropped = len(self.rows) - len(keep)
|
elif row.get("dedupe_veto"):
|
||||||
if dropped:
|
if new_title and row["title_raw"] != new_title:
|
||||||
|
row["title_raw"] = new_title
|
||||||
|
renamed += 1
|
||||||
|
keep.append(row)
|
||||||
|
else:
|
||||||
|
dropped += 1
|
||||||
|
if dropped or renamed:
|
||||||
self.rows = keep
|
self.rows = keep
|
||||||
self._save()
|
self._write_rows()
|
||||||
return dropped
|
return dropped
|
||||||
|
|
||||||
def veto_merge(self, row: dict) -> None:
|
def veto_merge(self, row: dict) -> None:
|
||||||
|
|||||||
@@ -310,6 +310,13 @@ button.danger { color: var(--stop-ink); border-color: var(--stop); background: #
|
|||||||
text-underline-offset: 2px;
|
text-underline-offset: 2px;
|
||||||
}
|
}
|
||||||
.catalog a:hover { color: var(--accent-ink); }
|
.catalog a:hover { color: var(--accent-ink); }
|
||||||
|
.catalog td.actions { text-align: right; white-space: nowrap; }
|
||||||
|
.catalog td.actions button {
|
||||||
|
font: inherit; font-size: .78rem; border: var(--line); background: #fff;
|
||||||
|
border-radius: var(--radius); padding: .2rem .6rem; cursor: pointer;
|
||||||
|
box-shadow: none;
|
||||||
|
}
|
||||||
|
.catalog td.actions button:hover { background: var(--board); }
|
||||||
.catalog tr.editrow td { background: #fff; border-top: none; padding: .2rem .5rem .7rem; }
|
.catalog tr.editrow td { background: #fff; border-top: none; padding: .2rem .5rem .7rem; }
|
||||||
.editform { display: flex; gap: .7rem; align-items: end; flex-wrap: wrap; }
|
.editform { display: flex; gap: .7rem; align-items: end; flex-wrap: wrap; }
|
||||||
.editform label {
|
.editform label {
|
||||||
@@ -319,10 +326,9 @@ button.danger { color: var(--stop-ink); border-color: var(--stop); background: #
|
|||||||
}
|
}
|
||||||
.editform input {
|
.editform input {
|
||||||
font: inherit; font-size: .85rem; color: var(--ink);
|
font: inherit; font-size: .85rem; color: var(--ink);
|
||||||
border: 1px solid var(--board-edge); border-radius: var(--radius);
|
border: 2px solid var(--board-edge); border-radius: var(--radius);
|
||||||
padding: .25rem .45rem; background: var(--board);
|
padding: .25rem .45rem; background: #fff;
|
||||||
}
|
}
|
||||||
.editform input:focus { outline: 2px solid var(--accent); outline-offset: 1px; }
|
|
||||||
.editactions { display: flex; gap: .5rem; }
|
.editactions { display: flex; gap: .5rem; }
|
||||||
.edithint { font-size: .78rem; color: var(--ink-soft); align-self: center; }
|
.edithint { font-size: .78rem; color: var(--ink-soft); align-self: center; }
|
||||||
.chip {
|
.chip {
|
||||||
|
|||||||
@@ -52,7 +52,7 @@ function render() {
|
|||||||
<td class="meta">${c.photos.map(p =>
|
<td class="meta">${c.photos.map(p =>
|
||||||
`<a href="/photos/view/${encodeURIComponent(p)}">${esc(p)}</a>`
|
`<a href="/photos/view/${encodeURIComponent(p)}">${esc(p)}</a>`
|
||||||
).join(", ")}</td>
|
).join(", ")}</td>
|
||||||
<td class="rowactions">${c.can_split
|
<td class="actions">${c.can_split
|
||||||
? `<button class="split" data-title="${esc(c.title_raw)}"
|
? `<button class="split" data-title="${esc(c.title_raw)}"
|
||||||
data-photos="${esc(c.photos.join(";"))}"
|
data-photos="${esc(c.photos.join(";"))}"
|
||||||
data-rowix="${c.row_ix ?? ""}"
|
data-rowix="${c.row_ix ?? ""}"
|
||||||
|
|||||||
+91
-36
@@ -39,6 +39,8 @@ from rich.console import Console
|
|||||||
from bggpipe.bgg_client import BGGClient, cached_paths
|
from bggpipe.bgg_client import BGGClient, cached_paths
|
||||||
from bggpipe.config import DEFAULT_REVIEW_PORT, Config
|
from bggpipe.config import DEFAULT_REVIEW_PORT, Config
|
||||||
from bggpipe.extract import (
|
from bggpipe.extract import (
|
||||||
|
is_split,
|
||||||
|
load_title_splits,
|
||||||
record_title_edit,
|
record_title_edit,
|
||||||
record_title_split,
|
record_title_split,
|
||||||
replay_titles,
|
replay_titles,
|
||||||
@@ -49,6 +51,8 @@ from bggpipe.models import (
|
|||||||
CONFIDENT_VERSION_STATUSES,
|
CONFIDENT_VERSION_STATUSES,
|
||||||
RECOGNIZED_MATCH_STATUSES,
|
RECOGNIZED_MATCH_STATUSES,
|
||||||
)
|
)
|
||||||
|
from bggpipe.normalize import normalize_title
|
||||||
|
from bggpipe.resolve import TitleEntry
|
||||||
from bggpipe.review import ReviewSession
|
from bggpipe.review import ReviewSession
|
||||||
|
|
||||||
|
|
||||||
@@ -155,7 +159,8 @@ class SplitBody(BaseModel):
|
|||||||
|
|
||||||
class EditBody(BaseModel):
|
class EditBody(BaseModel):
|
||||||
"""A human correction to an extracted read. None = leave that field
|
"""A human correction to an extracted read. None = leave that field
|
||||||
alone; empty string = clear it."""
|
alone; for the cue fields, empty string = clear it (a corrected title
|
||||||
|
may not be empty)."""
|
||||||
|
|
||||||
title_raw: str
|
title_raw: str
|
||||||
source_photos: str = ""
|
source_photos: str = ""
|
||||||
@@ -322,6 +327,18 @@ def create_app(
|
|||||||
app_warnings.append(note)
|
app_warnings.append(note)
|
||||||
return {}
|
return {}
|
||||||
|
|
||||||
|
def _find_entry(title_raw: str, source_photos: str) -> TitleEntry | None:
|
||||||
|
photos = [p for p in source_photos.split(";") if p]
|
||||||
|
return next(
|
||||||
|
(
|
||||||
|
e
|
||||||
|
for e in session.titles
|
||||||
|
if e.title_raw == title_raw
|
||||||
|
and (not photos or list(e.source_photos) == photos)
|
||||||
|
),
|
||||||
|
None,
|
||||||
|
)
|
||||||
|
|
||||||
def find_row(title_raw: str, source_photos: str, row_ix: int | None = None) -> dict:
|
def find_row(title_raw: str, source_photos: str, row_ix: int | None = None) -> dict:
|
||||||
freshen()
|
freshen()
|
||||||
row = session.find_row(title_raw, source_photos, row_ix)
|
row = session.find_row(title_raw, source_photos, row_ix)
|
||||||
@@ -380,6 +397,9 @@ def create_app(
|
|||||||
catalog = []
|
catalog = []
|
||||||
|
|
||||||
def catalog_line(entry, row) -> dict:
|
def catalog_line(entry, row) -> dict:
|
||||||
|
cue_entry = entry or (
|
||||||
|
_find_entry(row["title_raw"], row["source_photos"]) if row else None
|
||||||
|
)
|
||||||
return {
|
return {
|
||||||
"title_raw": entry.title_raw if entry else row["title_raw"],
|
"title_raw": entry.title_raw if entry else row["title_raw"],
|
||||||
"confidence": entry.confidence if entry else "",
|
"confidence": entry.confidence if entry else "",
|
||||||
@@ -396,10 +416,10 @@ def create_app(
|
|||||||
"merged_into": row.get("merged_into", "") if row else "",
|
"merged_into": row.get("merged_into", "") if row else "",
|
||||||
"row_ix": _ix_of(session.rows, row) if row else None,
|
"row_ix": _ix_of(session.rows, row) if row else None,
|
||||||
"cues": {
|
"cues": {
|
||||||
"publisher": entry.publisher_hint if entry else "",
|
"publisher": cue_entry.publisher_hint if cue_entry else "",
|
||||||
"edition": entry.edition_hint if entry else "",
|
"edition": cue_entry.edition_hint if cue_entry else "",
|
||||||
"year": entry.year_hint if entry else None,
|
"year": cue_entry.year_hint if cue_entry else None,
|
||||||
"language": entry.language_hint if entry else "",
|
"language": cue_entry.language_hint if cue_entry else "",
|
||||||
},
|
},
|
||||||
"can_split": bool(
|
"can_split": bool(
|
||||||
row
|
row
|
||||||
@@ -417,8 +437,10 @@ def create_app(
|
|||||||
same_title = rows_by_title.get(entry.title_raw, [])
|
same_title = rows_by_title.get(entry.title_raw, [])
|
||||||
ix = title_seen[entry.title_raw]
|
ix = title_seen[entry.title_raw]
|
||||||
title_seen[entry.title_raw] += 1
|
title_seen[entry.title_raw] += 1
|
||||||
# positional pairing, same rule as run_resolve: the ix-th entry
|
# positional pairing: the ix-th entry of a title reports the
|
||||||
# of a title reports the ix-th row of that title
|
# ix-th row of that title (run_resolve pairs photo-overlap
|
||||||
|
# first, but split entries and split rows both keep sorted
|
||||||
|
# photo order, so the ordinals line up)
|
||||||
row = same_title[ix] if ix < len(same_title) else None
|
row = same_title[ix] if ix < len(same_title) else None
|
||||||
# a split row's photo set is narrower than its entry's — show
|
# a split row's photo set is narrower than its entry's — show
|
||||||
# the row's own photos for split copies
|
# the row's own photos for split copies
|
||||||
@@ -722,18 +744,6 @@ def create_app(
|
|||||||
raise HTTPException(400, f"unknown action {body.action!r}")
|
raise HTTPException(400, f"unknown action {body.action!r}")
|
||||||
return state()
|
return state()
|
||||||
|
|
||||||
def _find_entry(title_raw: str, source_photos: str):
|
|
||||||
photos = [p for p in source_photos.split(";") if p]
|
|
||||||
return next(
|
|
||||||
(
|
|
||||||
e
|
|
||||||
for e in session.titles
|
|
||||||
if e.title_raw == title_raw
|
|
||||||
and (not photos or list(e.source_photos) == photos)
|
|
||||||
),
|
|
||||||
None,
|
|
||||||
)
|
|
||||||
|
|
||||||
@app.post("/api/split")
|
@app.post("/api/split")
|
||||||
def api_split(body: SplitBody) -> dict:
|
def api_split(body: SplitBody) -> dict:
|
||||||
with lock:
|
with lock:
|
||||||
@@ -743,9 +753,18 @@ def create_app(
|
|||||||
row = session.find_row(body.title_raw, body.source_photos, body.row_ix)
|
row = session.find_row(body.title_raw, body.source_photos, body.row_ix)
|
||||||
if row is not None:
|
if row is not None:
|
||||||
try:
|
try:
|
||||||
session.split_row(row)
|
copies = session.split_row(row)
|
||||||
except ValueError as err:
|
except ValueError as err:
|
||||||
raise HTTPException(400, str(err)) from err
|
raise HTTPException(400, str(err)) from err
|
||||||
|
if not copies:
|
||||||
|
# split_row adopted nothing: matches.csv was rewritten
|
||||||
|
# underneath — recording the store now would report a
|
||||||
|
# success that didn't happen
|
||||||
|
raise HTTPException(
|
||||||
|
409,
|
||||||
|
"matches.csv changed while you split — nothing "
|
||||||
|
"changed; retry from the refreshed page",
|
||||||
|
)
|
||||||
else:
|
else:
|
||||||
# no matches row yet — the title is still awaiting resolve;
|
# no matches row yet — the title is still awaiting resolve;
|
||||||
# splitting is purely a titles.json (extraction) decision
|
# splitting is purely a titles.json (extraction) decision
|
||||||
@@ -761,7 +780,11 @@ def create_app(
|
|||||||
# persist the decision so every future extract rebuild and
|
# persist the decision so every future extract rebuild and
|
||||||
# resolve dedupe keeps the copies apart, then re-derive
|
# resolve dedupe keeps the copies apart, then re-derive
|
||||||
# titles.json so each copy carries its own photo's cues
|
# titles.json so each copy carries its own photo's cues
|
||||||
record_title_split(cfg.title_splits_path, body.title_raw)
|
record_title_split(
|
||||||
|
cfg.title_splits_path,
|
||||||
|
body.title_raw,
|
||||||
|
[p for p in body.source_photos.split(";") if p],
|
||||||
|
)
|
||||||
replay_titles(cfg)
|
replay_titles(cfg)
|
||||||
return state()
|
return state()
|
||||||
|
|
||||||
@@ -778,11 +801,34 @@ def create_app(
|
|||||||
)
|
)
|
||||||
record: dict = {"match": body.title_raw}
|
record: dict = {"match": body.title_raw}
|
||||||
photos = [p for p in body.source_photos.split(";") if p]
|
photos = [p for p in body.source_photos.split(";") if p]
|
||||||
same_title = [e for e in session.titles if e.title_raw == body.title_raw]
|
norm = normalize_title(body.title_raw)
|
||||||
if photos and len(same_title) > 1:
|
# scope by NORMALIZED title: stored edits apply by normalized
|
||||||
|
# match, so raw-equality counting here would let an edit bleed
|
||||||
|
# onto a differently-cased sighting of another edition
|
||||||
|
same_norm = [
|
||||||
|
e for e in session.titles if normalize_title(e.title_raw) == norm
|
||||||
|
]
|
||||||
|
if photos and len(same_norm) > 1:
|
||||||
# several copies/editions share this title: the fix targets
|
# several copies/editions share this title: the fix targets
|
||||||
# only the copy the human was looking at
|
# only the copy the human was looking at
|
||||||
record["photos"] = photos
|
record["photos"] = photos
|
||||||
|
if (
|
||||||
|
sum(
|
||||||
|
1
|
||||||
|
for e in same_norm
|
||||||
|
if list(e.source_photos) == list(entry.source_photos)
|
||||||
|
)
|
||||||
|
> 1
|
||||||
|
):
|
||||||
|
# two same-named editions seen in the same photo(s): the
|
||||||
|
# store can only target (title, photos), so an edit would
|
||||||
|
# hit both — refuse rather than corrupt the sibling
|
||||||
|
raise HTTPException(
|
||||||
|
409,
|
||||||
|
"two entries share this exact title and photo set — an "
|
||||||
|
"edit cannot target just one; resolve or reject one of "
|
||||||
|
"them first",
|
||||||
|
)
|
||||||
if body.title_new is not None:
|
if body.title_new is not None:
|
||||||
corrected = body.title_new.strip()
|
corrected = body.title_new.strip()
|
||||||
if not corrected:
|
if not corrected:
|
||||||
@@ -801,22 +847,31 @@ def create_app(
|
|||||||
if year and not year.isdigit():
|
if year and not year.isdigit():
|
||||||
raise HTTPException(400, "year must be a number")
|
raise HTTPException(400, "year must be a number")
|
||||||
record["year_hint"] = int(year) if year else None
|
record["year_hint"] = int(year) if year else None
|
||||||
if not any(
|
if not (record.keys() - {"match", "photos"}):
|
||||||
key in record
|
|
||||||
for key in (
|
|
||||||
"title_raw",
|
|
||||||
"publisher_hint",
|
|
||||||
"edition_hint",
|
|
||||||
"language_hint",
|
|
||||||
"year_hint",
|
|
||||||
)
|
|
||||||
):
|
|
||||||
raise HTTPException(400, "nothing to change")
|
raise HTTPException(400, "nothing to change")
|
||||||
|
# write order is crash-safety: drop the stale rows first (worst
|
||||||
|
# case on a crash: resolve recreates them from the uncorrected
|
||||||
|
# entry), then the durable record (replayed by every future
|
||||||
|
# rebuild), then the titles.json replay
|
||||||
|
session.drop_rows(
|
||||||
|
body.title_raw,
|
||||||
|
record.get("photos"),
|
||||||
|
new_title=record.get("title_raw"),
|
||||||
|
)
|
||||||
|
if "title_raw" in record and is_split(
|
||||||
|
norm,
|
||||||
|
entry.source_photos,
|
||||||
|
load_title_splits(cfg.title_splits_path),
|
||||||
|
):
|
||||||
|
# a renamed split copy must stay under split protection —
|
||||||
|
# the old record keys the OLD normalized title
|
||||||
|
record_title_split(
|
||||||
|
cfg.title_splits_path,
|
||||||
|
record["title_raw"],
|
||||||
|
list(entry.source_photos),
|
||||||
|
)
|
||||||
record_title_edit(cfg.title_edits_path, record)
|
record_title_edit(cfg.title_edits_path, record)
|
||||||
replay_titles(cfg)
|
replay_titles(cfg)
|
||||||
# rows matched against the uncorrected read are stale: drop
|
|
||||||
# them so the next resolve re-queries with the fix
|
|
||||||
session.drop_rows(body.title_raw, photos if "photos" in record else None)
|
|
||||||
return state()
|
return state()
|
||||||
|
|
||||||
@app.post("/api/veto-merge")
|
@app.post("/api/veto-merge")
|
||||||
|
|||||||
+49
-4
@@ -177,11 +177,34 @@ def test_dedupe_conflicting_years_stay_separate():
|
|||||||
def test_dedupe_split_titles_never_merge():
|
def test_dedupe_split_titles_never_merge():
|
||||||
deduped = dedupe_entries(
|
deduped = dedupe_entries(
|
||||||
[_entry("Wiz-War", "a.jpg"), _entry("Wiz-War", "b.jpg")],
|
[_entry("Wiz-War", "a.jpg"), _entry("Wiz-War", "b.jpg")],
|
||||||
split_titles={"wiz war"},
|
splits=[{"norm": "wiz war", "photos": None}],
|
||||||
)
|
)
|
||||||
assert len(deduped) == 2 # the human said: separate physical copies
|
assert len(deduped) == 2 # the human said: separate physical copies
|
||||||
|
|
||||||
|
|
||||||
|
def test_photo_scoped_split_spares_other_editions():
|
||||||
|
# splitting the copies seen in a/b must not force-split a same-named
|
||||||
|
# different edition (c/d, kept separate by its conflicting cue)
|
||||||
|
entries = [
|
||||||
|
_entry("Carcassonne", "a.jpg", publisher_hint="Rio Grande"),
|
||||||
|
_entry("Carcassonne", "b.jpg", publisher_hint="Rio Grande"),
|
||||||
|
_entry("Carcassonne", "c.jpg", publisher_hint="Z-Man"),
|
||||||
|
_entry("Carcassonne", "d.jpg", publisher_hint="Z-Man"),
|
||||||
|
]
|
||||||
|
splits = [{"norm": "carcassonne", "photos": {"a.jpg", "b.jpg"}}]
|
||||||
|
deduped = dedupe_entries(entries, splits)
|
||||||
|
photo_sets = [e["source_photos"] for e in deduped]
|
||||||
|
assert ["a.jpg"] in photo_sets and ["b.jpg"] in photo_sets # split copies
|
||||||
|
assert ["c.jpg", "d.jpg"] in photo_sets # other edition still dedupes
|
||||||
|
|
||||||
|
|
||||||
|
def test_corrupt_store_fails_loud_with_filename(tmp_path):
|
||||||
|
path = tmp_path / "title_splits.json"
|
||||||
|
path.write_text("<<<<<<< merge conflict")
|
||||||
|
with pytest.raises(ValueError, match="title_splits.json"):
|
||||||
|
load_title_splits(path)
|
||||||
|
|
||||||
|
|
||||||
def test_edits_fix_misreads_before_dedupe():
|
def test_edits_fix_misreads_before_dedupe():
|
||||||
# a corrected misspelling merges with the correctly-read sighting
|
# a corrected misspelling merges with the correctly-read sighting
|
||||||
edits = [{"match": "Hebarceos", "title_raw": "Herbaceous"}]
|
edits = [{"match": "Hebarceos", "title_raw": "Herbaceous"}]
|
||||||
@@ -219,12 +242,16 @@ def test_stores_roundtrip_and_replay_from_raw(tmp_path):
|
|||||||
(raw / f"{photo}.json").write_text(
|
(raw / f"{photo}.json").write_text(
|
||||||
json.dumps({"titles": [_entry("Wiz-War", photo)], "unidentified": []})
|
json.dumps({"titles": [_entry("Wiz-War", photo)], "unidentified": []})
|
||||||
)
|
)
|
||||||
record_title_split(cfg.title_splits_path, "Wiz-War")
|
record_title_split(cfg.title_splits_path, "Wiz-War", ["a.jpg", "b.jpg"])
|
||||||
record_title_split(cfg.title_splits_path, "wiz war") # dupe, normalized away
|
record_title_split(cfg.title_splits_path, "wiz war", ["a.jpg"]) # covered: no dupe
|
||||||
record_title_edit(
|
record_title_edit(
|
||||||
cfg.title_edits_path,
|
cfg.title_edits_path,
|
||||||
{"match": "Wiz-War", "photos": ["a.jpg"], "edition_hint": "7th Edition"},
|
{"match": "Wiz-War", "photos": ["a.jpg"], "edition_hint": "7th Edition"},
|
||||||
)
|
)
|
||||||
|
record_title_edit( # identical retry must not double-record
|
||||||
|
cfg.title_edits_path,
|
||||||
|
{"match": "Wiz-War", "photos": ["a.jpg"], "edition_hint": "7th Edition"},
|
||||||
|
)
|
||||||
replay_titles(cfg)
|
replay_titles(cfg)
|
||||||
titles = json.loads(cfg.titles_path.read_text())
|
titles = json.loads(cfg.titles_path.read_text())
|
||||||
assert [e["source_photos"] for e in titles] == [["a.jpg"], ["b.jpg"]]
|
assert [e["source_photos"] for e in titles] == [["a.jpg"], ["b.jpg"]]
|
||||||
@@ -242,13 +269,31 @@ def test_stores_roundtrip_and_replay_from_raw(tmp_path):
|
|||||||
assert len(json.loads(cfg.titles_path.read_text())) == 2
|
assert len(json.loads(cfg.titles_path.read_text())) == 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_replay_ignores_partial_raw_caches(tmp_path):
|
||||||
|
# fresh-clone shape: committed titles.json spans two photos, but only
|
||||||
|
# one raw cache file exists (raw is gitignored) — replay must not
|
||||||
|
# rebuild from the partial raws and truncate the catalog
|
||||||
|
cfg = Config(data_dir=tmp_path / "data", photos_dir=tmp_path / "photos")
|
||||||
|
raw = cfg.extract_raw_dir
|
||||||
|
raw.mkdir(parents=True)
|
||||||
|
(raw / "a.jpg.json").write_text(
|
||||||
|
json.dumps({"titles": [_entry("Catan", "a.jpg")], "unidentified": []})
|
||||||
|
)
|
||||||
|
cfg.titles_path.write_text(
|
||||||
|
json.dumps([_entry("Catan", "a.jpg"), _entry("Wingspan", "b.jpg")])
|
||||||
|
)
|
||||||
|
replay_titles(cfg)
|
||||||
|
titles = {e["title_raw"] for e in json.loads(cfg.titles_path.read_text())}
|
||||||
|
assert titles == {"Catan", "Wingspan"}
|
||||||
|
|
||||||
|
|
||||||
def test_replay_without_raw_caches_explodes_merged_entries(tmp_path):
|
def test_replay_without_raw_caches_explodes_merged_entries(tmp_path):
|
||||||
cfg = Config(data_dir=tmp_path / "data", photos_dir=tmp_path / "photos")
|
cfg = Config(data_dir=tmp_path / "data", photos_dir=tmp_path / "photos")
|
||||||
cfg.data_dir.mkdir(parents=True)
|
cfg.data_dir.mkdir(parents=True)
|
||||||
merged = _entry("Wiz-War", "a.jpg")
|
merged = _entry("Wiz-War", "a.jpg")
|
||||||
merged["source_photos"] = ["a.jpg", "b.jpg", "c.jpg"]
|
merged["source_photos"] = ["a.jpg", "b.jpg", "c.jpg"]
|
||||||
cfg.titles_path.write_text(json.dumps([merged, _entry("Catan", "a.jpg")]))
|
cfg.titles_path.write_text(json.dumps([merged, _entry("Catan", "a.jpg")]))
|
||||||
record_title_split(cfg.title_splits_path, "Wiz-War")
|
record_title_split(cfg.title_splits_path, "Wiz-War", ["a.jpg", "b.jpg", "c.jpg"])
|
||||||
replay_titles(cfg)
|
replay_titles(cfg)
|
||||||
titles = json.loads(cfg.titles_path.read_text())
|
titles = json.loads(cfg.titles_path.read_text())
|
||||||
by_title = {}
|
by_title = {}
|
||||||
|
|||||||
@@ -444,7 +444,7 @@ def test_dedupe_matches_skips_split_titles():
|
|||||||
_mrow("Wiz-War", "94", "a.jpg"),
|
_mrow("Wiz-War", "94", "a.jpg"),
|
||||||
_mrow("Wiz-War", "94", "b.jpg"),
|
_mrow("Wiz-War", "94", "b.jpg"),
|
||||||
]
|
]
|
||||||
assert dedupe_matches(rows, [], split_titles={"wiz war"}) == []
|
assert dedupe_matches(rows, [], splits=[{"norm": "wiz war", "photos": None}]) == []
|
||||||
assert all(not r["merged_into"] for r in rows)
|
assert all(not r["merged_into"] for r in rows)
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
+169
-4
@@ -477,8 +477,10 @@ def test_rowless_multiphoto_title_is_splittable(tmp_path):
|
|||||||
copies = [c for c in res.json()["catalog"] if c["title_raw"] == "Wiz-War"]
|
copies = [c for c in res.json()["catalog"] if c["title_raw"] == "Wiz-War"]
|
||||||
assert [c["photos"] for c in copies] == [["shelf.jpg"], ["shelf2.jpg"]]
|
assert [c["photos"] for c in copies] == [["shelf.jpg"], ["shelf2.jpg"]]
|
||||||
assert all(c["can_split"] is False for c in copies)
|
assert all(c["can_split"] is False for c in copies)
|
||||||
# the decision persisted: the store holds it and titles.json is split
|
# the decision persisted: the store holds it (photo-scoped) and
|
||||||
assert "Wiz-War" in json.loads(cfg.title_splits_path.read_text())
|
# titles.json is split
|
||||||
|
(record,) = json.loads(cfg.title_splits_path.read_text())
|
||||||
|
assert record == {"title": "Wiz-War", "photos": ["shelf.jpg", "shelf2.jpg"]}
|
||||||
per_photo = [
|
per_photo = [
|
||||||
e["source_photos"]
|
e["source_photos"]
|
||||||
for e in json.loads(cfg.titles_path.read_text())
|
for e in json.loads(cfg.titles_path.read_text())
|
||||||
@@ -510,7 +512,8 @@ def test_edit_title_corrects_read_and_requeues_row(tmp_path):
|
|||||||
"title_raw": "Citadels",
|
"title_raw": "Citadels",
|
||||||
"source_photos": "shelf.jpg",
|
"source_photos": "shelf.jpg",
|
||||||
"title_new": "Citadels: Dark City",
|
"title_new": "Citadels: Dark City",
|
||||||
"edition_hint": None,
|
"edition": "2nd Edition",
|
||||||
|
"publisher": "",
|
||||||
"year": "2004",
|
"year": "2004",
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
@@ -524,10 +527,13 @@ def test_edit_title_corrects_read_and_requeues_row(tmp_path):
|
|||||||
if e["title_raw"] == "Citadels: Dark City"
|
if e["title_raw"] == "Citadels: Dark City"
|
||||||
)
|
)
|
||||||
assert corrected["year_hint"] == 2004
|
assert corrected["year_hint"] == 2004
|
||||||
assert corrected["publisher_hint"] == "Fantasy Flight" # untouched cue kept
|
assert corrected["edition_hint"] == "2nd Edition"
|
||||||
|
assert corrected["publisher_hint"] == "" # empty string cleared the cue
|
||||||
|
assert not corrected.get("language_hint") # untouched (was unset)
|
||||||
assert not any(r["title_raw"] == "Citadels" for r in read_matches(cfg.matches_path))
|
assert not any(r["title_raw"] == "Citadels" for r in read_matches(cfg.matches_path))
|
||||||
(record,) = json.loads(cfg.title_edits_path.read_text())
|
(record,) = json.loads(cfg.title_edits_path.read_text())
|
||||||
assert record["match"] == "Citadels"
|
assert record["match"] == "Citadels"
|
||||||
|
assert record["publisher_hint"] == ""
|
||||||
|
|
||||||
|
|
||||||
def test_edit_title_rejects_empty_and_noop(tmp_path):
|
def test_edit_title_rejects_empty_and_noop(tmp_path):
|
||||||
@@ -561,3 +567,162 @@ def test_edit_title_rejects_empty_and_noop(tmp_path):
|
|||||||
).status_code
|
).status_code
|
||||||
== 404
|
== 404
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_edit_preserves_vetoed_rows_and_renames_them(tmp_path):
|
||||||
|
cfg = make_cfg(tmp_path)
|
||||||
|
rows = read_matches(cfg.matches_path)
|
||||||
|
rows.append(
|
||||||
|
_row(
|
||||||
|
title_raw="Citadels",
|
||||||
|
match_status="approved",
|
||||||
|
bgg_id="478",
|
||||||
|
source_photos="shelf2.jpg",
|
||||||
|
dedupe_veto="1",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
write_matches(cfg.matches_path, rows)
|
||||||
|
web = TestClient(create_app(cfg, client=unauthorized_client(tmp_path)))
|
||||||
|
res = web.post(
|
||||||
|
"/api/edit-title",
|
||||||
|
json={
|
||||||
|
"title_raw": "Citadels",
|
||||||
|
"source_photos": "shelf.jpg",
|
||||||
|
"title_new": "Citadels (2016)",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
assert res.status_code == 200
|
||||||
|
after = read_matches(cfg.matches_path)
|
||||||
|
# the human-vetoed row survives the cull, renamed to follow the entry;
|
||||||
|
# the unvetoed ambiguous row is requeued (dropped)
|
||||||
|
veto = [r for r in after if r.get("dedupe_veto")]
|
||||||
|
assert len(veto) == 1
|
||||||
|
assert veto[0]["title_raw"] == "Citadels (2016)"
|
||||||
|
assert veto[0]["match_status"] == "approved"
|
||||||
|
assert not any(
|
||||||
|
r["title_raw"] == "Citadels" and not r.get("dedupe_veto") for r in after
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_drop_rows_photo_narrowing_spares_other_copy(tmp_path):
|
||||||
|
import io as _io
|
||||||
|
|
||||||
|
from rich.console import Console
|
||||||
|
|
||||||
|
from bggpipe.review import ReviewSession
|
||||||
|
|
||||||
|
cfg = make_cfg(tmp_path)
|
||||||
|
write_matches(
|
||||||
|
cfg.matches_path,
|
||||||
|
[
|
||||||
|
_row(title_raw="Wiz-War", match_status="auto", source_photos="a.jpg"),
|
||||||
|
_row(title_raw="Wiz-War", match_status="auto", source_photos="b.jpg"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
session = ReviewSession(
|
||||||
|
cfg,
|
||||||
|
console=Console(file=_io.StringIO()),
|
||||||
|
input_fn=lambda prompt: "",
|
||||||
|
client=unauthorized_client(tmp_path),
|
||||||
|
)
|
||||||
|
assert session.drop_rows("Wiz-War", ["a.jpg"]) == 1
|
||||||
|
(survivor,) = read_matches(cfg.matches_path)
|
||||||
|
assert survivor["source_photos"] == "b.jpg"
|
||||||
|
|
||||||
|
|
||||||
|
def test_split_and_edit_refuse_while_pipeline_rewrites(tmp_path):
|
||||||
|
import threading
|
||||||
|
|
||||||
|
from bggpipe.jobs import JobRunner
|
||||||
|
|
||||||
|
cfg = make_cfg(tmp_path)
|
||||||
|
_rowless_wizwar(cfg)
|
||||||
|
release = threading.Event()
|
||||||
|
started = threading.Event()
|
||||||
|
|
||||||
|
def blocking_extract():
|
||||||
|
started.set()
|
||||||
|
release.wait(timeout=5)
|
||||||
|
|
||||||
|
jobs = JobRunner()
|
||||||
|
web = TestClient(
|
||||||
|
create_app(
|
||||||
|
cfg,
|
||||||
|
client=unauthorized_client(tmp_path),
|
||||||
|
stages={"extract": blocking_extract},
|
||||||
|
jobs=jobs,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
assert web.post("/api/run/extract").status_code == 200
|
||||||
|
assert started.wait(timeout=5)
|
||||||
|
try:
|
||||||
|
split = web.post(
|
||||||
|
"/api/split",
|
||||||
|
json={
|
||||||
|
"title_raw": "Wiz-War",
|
||||||
|
"source_photos": "shelf.jpg;shelf2.jpg",
|
||||||
|
"row_ix": None,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
edit = web.post(
|
||||||
|
"/api/edit-title",
|
||||||
|
json={
|
||||||
|
"title_raw": "Wiz-War",
|
||||||
|
"source_photos": "shelf.jpg;shelf2.jpg",
|
||||||
|
"title_new": "Wiz-War!",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
finally:
|
||||||
|
release.set()
|
||||||
|
assert split.status_code == 409
|
||||||
|
assert edit.status_code == 409
|
||||||
|
# neither curation store was written under the in-flight rewrite
|
||||||
|
assert not cfg.title_splits_path.exists()
|
||||||
|
assert not cfg.title_edits_path.exists()
|
||||||
|
|
||||||
|
|
||||||
|
def test_renaming_a_split_copy_keeps_split_protection(tmp_path):
|
||||||
|
cfg = make_cfg(tmp_path)
|
||||||
|
_rowless_wizwar(cfg)
|
||||||
|
web = TestClient(create_app(cfg, client=unauthorized_client(tmp_path)))
|
||||||
|
assert (
|
||||||
|
web.post(
|
||||||
|
"/api/split",
|
||||||
|
json={
|
||||||
|
"title_raw": "Wiz-War",
|
||||||
|
"source_photos": "shelf.jpg;shelf2.jpg",
|
||||||
|
"row_ix": None,
|
||||||
|
},
|
||||||
|
).status_code
|
||||||
|
== 200
|
||||||
|
)
|
||||||
|
# rename BOTH copies to the same corrected title, one at a time — the
|
||||||
|
# store must keep protecting them or the rebuild re-merges the copies
|
||||||
|
for photo in ("shelf.jpg", "shelf2.jpg"):
|
||||||
|
assert (
|
||||||
|
web.post(
|
||||||
|
"/api/edit-title",
|
||||||
|
json={
|
||||||
|
"title_raw": "Wiz-War",
|
||||||
|
"source_photos": photo,
|
||||||
|
"title_new": "Wiz War 2000",
|
||||||
|
},
|
||||||
|
).status_code
|
||||||
|
== 200
|
||||||
|
)
|
||||||
|
per_photo = [
|
||||||
|
e["source_photos"]
|
||||||
|
for e in json.loads(cfg.titles_path.read_text())
|
||||||
|
if e["title_raw"] == "Wiz War 2000"
|
||||||
|
]
|
||||||
|
assert per_photo == [["shelf.jpg"], ["shelf2.jpg"]] # still two copies
|
||||||
|
|
||||||
|
|
||||||
|
def test_single_photo_rowless_entry_is_not_splittable(tmp_path):
|
||||||
|
web, _ = make_client(tmp_path)
|
||||||
|
line = next(
|
||||||
|
c
|
||||||
|
for c in web.get("/api/state").json()["catalog"]
|
||||||
|
if c["title_raw"] == "Fresh Off The Shelf"
|
||||||
|
)
|
||||||
|
assert line["can_split"] is False
|
||||||
|
|||||||
Reference in New Issue
Block a user