Commit record

Spot-check auto-caption quality across eras; add assessments

Committed

a03c8e79bbd679a6fb65e554a3d2e8b910aef43d

← All changes
Spot-check auto-caption quality across eras; add assessments New stage-2b tool scripts/transcribe_video.py (direct download with decodo-proxy fallback, S3 audio archive to s3://joedistilled/media/audio/, butler transcription, full videos.db registration) and scripts/compare_transcripts.py (token-alignment metrics between transcript versions). Schema v2 adds the assessments table recording head-to-head comparisons (metrics / llm-judge / human) so canonical promotions cite evidence. Butler re-transcribed 7 videos (4 from 2015-2016, 3 from 2024-2026; 8 pairs assessed counting the channel's first video). Verdict: no systematic era problem -- 7 of 8 pairs agree at 0.97-1.00 similarity. Old-era captions garble occasional key terms ("oh man" for OMAD, "ay hoffler" for Ori Hofmekler); recent ones are near-verbatim. Canonical flipped to butler for the 3 videos where it was judged strictly better (plus 4gcAjbAlcxk earlier); ties keep youtube-auto. Audio for all 7 archived in S3 and registered in media.