Introduce videos.db as the corpus source of truth
Normalized SQLite database (repo root, committed) replacing the TSV
era: videos (own stable ids, kind, exact UTC dates), transcripts (one
row per version -- youtube-auto vs butler -- with sha256, capture time,
STT model, proxy use, no-speech flag) with a canonical pointer per
video, text_files recording which transcript version generated each
markdown rendering, media for audio archived to
s3://joedistilled/media/audio/, surveys/survey_entries for channel
snapshots, and an append-only log of operations and failures. Schema in
scripts/videos_db_schema.sql; one-time migration in
scripts/migrate_tsv_to_db.py seeds today's backfill history.
Also: adopt the butler transcription of the channel's first video
4gcAjbAlcxk as canonical (its YouTube auto-captions dropped ~1/3 of the
content, including Joe's peak weight and the origin story); make
json3_to_text.py database-aware (renders canonical transcripts, records
provenance); retire catalog.tsv/index.tsv; update methodology and
transcripts README. Sub-minute durations now render as 0:SS.