Skip to content

Transcription

Built in, on the user's own machine, with no third-party service. Four stages, each replaceable by a plugin that declares the same activity.

1. Decode and detect speech

Find where someone is talking, so silence is not transcribed. On a two-hour interview with long pauses this is most of the saving.

2. Recognise

Whisper, at a model size the user chooses against their hardware. Podlibre suggests one based on what the machine can do and says plainly what the trade is: smaller is faster and wronger.

3. Attribute

Cluster the voices into speakers, which the user then names once per episode. Diarisation is the stage that turns a wall of text into something a reader can follow, and it is also the stage with the most variable quality — crosstalk is genuinely hard.

4. Correct

Ordered pattern rules from the user's dictionary. No model is involved. A podcaster who says Castopod forty times a season should fix the spelling once, for good, and watch it apply everywhere.

A rule is a pattern, a replacement, and a scope. Rules live in the show file, so they travel with the show, and a rule can be global or belong to one show. Suggestions are shown and never applied silently: the Edit screen highlights each one, Tab moves between them, and Accept all exists for when you trust your own dictionary.

What the user sees

Transcription is a task like any other: progress, an estimate, and a Stop button. It runs in a worker process, so a model that wedges can be killed without touching the window.