Files
wizwar6e/.claude/skills/wizwar-pulse/SKILL.md
T
Eric WagonerandClaude Fable 5.1 cf7920e159 The rollup counts the tags on posted links, and the pulse has a right-now
Reddit and BoardGameGeek send no referrer, so a visitor from either
looks direct. A ?ref=<name> on the posted link says where they came
from; the rollup counts those (one per address per day, Facebook's
fbclid as its own) beside the referrers browsers do send. The pulse
gains a right-now recipe: live sockets, rooms touched in the last
hour, and today's partial row.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm2auWk6RP71CjaAb4FMoG
2026-09-03 12:06:12 -04:00

4.8 KiB
Raw Blame History

name, description
name description
wizwar-pulse The weekly Wiz-War operations pulse — droplet health, error and traffic signals, backup status, and who has been playing. Use when Eric asks how the server is doing, wants an ops check, or says "run the pulse".

The Wiz-War ops pulse

One invocation, one screen: how the table is doing. Gather everything, then report in five short sections. Lead with anything ALARMING (service down, memory near the cap, failed backup, new Sentry issues); if all is quiet, say so in one line before the details.

Gather

  1. Box health (one ssh to root@104.236.96.198):
    • systemctl is-active wizwar caddy — both must be active
    • systemctl status wizwar Memory line — the cap is 700M; worry past ~500M steady-state
    • free -m, df -h /, uptime — swap in heavy use, disk past ~70%, or load near 1.0 (single vCPU) are the resize triggers
    • journalctl -u wizwar --since "-7 days" | grep -ci "unhandled protocol error" — anything nonzero deserves a look (Sentry has the stack traces)
  2. Sentry (MCP connector, org locallygrownnet, project wizwar):
    • search_issues for is:unresolved — new issues are the headline
    • get_monitor_details for cron monitor slug new-monitor (named wizwar-backup) — last check-in must be ok and recent; the uptime monitor alerts on its own but note any incidents
    • New-project caveat: results can lag ingest by minutes; the connector reads fine but 403s on create operations
  3. Backups: tail -3 /var/log/wizwar/backup.log (or its rotation) for the last run, plus the cron monitor above for missed runs
  4. Trends (/var/lib/wizwar/rollup.jsonl, one JSON line per UTC day written at 00:10 by /usr/local/bin/wizwar-rollup.sh; counts only, no addresses): the last 7 rows give requests, human vs bot addresses, socket connects, path mix (/, /clips, /watch, /ws), external referrers (reddit.com, boardgamegeek.com after the 2026-09-03 announcement), new rooms, human seats and names, reports/replies, and vitals (memMB, memPeakMB, diskPct, load1, protocolErrors, serviceStarts, ledgerMB). Report day-over-day direction, not raw rows. Its Sentry cron monitor is slug wizwar-rollup (a missed check-in there means the rollup itself is broken). Re-roll a day by hand: wizwar-rollup.sh YYYY-MM-DD. campaigns counts visitors (one per address per day) whose link carried ?ref=<name> — the tags on Eric's posts (reddit, bgg, …) — or a Facebook fbclid; referrers alone can't attribute Reddit or BGG, which send none. Right now (on request, or when the announcement is fresh): live sockets ss -tn state established '( sport = :8787 )' | tail -n +2 | wc -l, rooms touched in the last hour find /var/lib/wizwar/rooms -name '*.jsonl' -mmin -60 | wc -l, and today so far wizwar-rollup.sh $(date -u +%F) (a partial row, replaced again at 00:10; the vitals in it are the moment's).
  5. The table itself:
    • Room count and growth: ls /var/lib/wizwar/rooms/*.jsonl | wc -l
    • New human names this week: room ledgers' join lines minus bots (Automaton*) and known names, on rooms created in the last 7 days (meta line createdAt)
    • Unanswered feedback count from /var/lib/wizwar/feedback.jsonl — if any, offer to run the reports desk (/wizwar-reports)
    • Access-log pulse: wc -l /var/lib/caddy/access.log and a count of distinct client_ips in the last day, for the growth line

Report

Six sections, tight: Verdict (one line), Box, Errors & monitors, Backups, Trends (the week's direction from the rollup), The table (games, new faces, feedback). Numbers with their thresholds, not raw dumps. End with any recommended action, or "nothing needs you."

Standing knowledge

  • The service restarts cleanly (ledgers replay); restarts in the journal during deploys are normal, not incidents.
  • Room memory is evicted after idle (finished 30 min, others 24 h); boot-restores load everything then shed, so post-deploy memory spikes settle within the hour.
  • Unattended-upgrades reboots the box at 09:30 UTC when a kernel patch requires it; a reboot there is maintenance, not an outage.
  • Caddy access logs live at /var/lib/caddy/access.log (self-rotating, 10MiB × 30); the systemd sandbox denies /var/log/caddy.
  • Per-address limits (2026-09-03): 12 new rooms and 6 reports per address per hour, in packages/server/src/ratelimit.ts. A player who hits one sees "try again in an hour"; a pulse showing many refused creates from one address is a script, not a friend.
  • Restore drill (2026-09-01): a Spaces snapshot booted 56/56 rooms clean. The path: rclone copy spaces:kestrel-wizwar-backups/wizwar/ snapshots/<date> ... on the droplet, then run a local server with WIZWAR_DATA_DIR pointed at its rooms dir. Re-drill quarterly-ish.