Transcription¶
Built in, on the user's own machine, with no third-party service. Four stages, each replaceable by a plugin that declares the same activity.
1. Decode and detect speech¶
Find where someone is talking, so silence is not transcribed. On a two-hour interview with long pauses this is most of the saving.
2. Recognise¶
Whisper, at a model size the user chooses against their hardware. Podlibre suggests one based on what the machine can do and says plainly what the trade is: smaller is faster and wronger.
3. Attribute¶
Cluster the voices into speakers, which the user then names once per episode. Diarisation is the stage that turns a wall of text into something a reader can follow, and it is also the stage with the most variable quality — crosstalk is genuinely hard.
4. Correct¶
Ordered pattern rules from the user's dictionary. No model is involved. A podcaster who says Castopod forty times a season should fix the spelling once, for good, and watch it apply everywhere.
A rule is a pattern, a replacement, and a scope. Rules live in the show file, so they travel with
the show, and a rule can be global or belong to one show. Suggestions are shown and never applied
silently: the Edit screen highlights each one, Tab moves between them, and Accept all exists
for when you trust your own dictionary.
What the user sees¶
Transcription is a task like any other: progress, an estimate, and a Stop button. It runs in a worker process, so a model that wedges can be killed without touching the window.