Skip to content

Transcribe Videos

Fast Video Cataloger can generate text transcripts from the audio in your videos, making spoken content searchable and browsable. There are two ways to get transcripts into your catalog: the built-in speech-to-text, or importing existing SRT subtitle files.

Native Speech-to-Text Transcription

To enable automatic transcription during indexing, open Preferences, go to the Video Indexer tab, and check Enable speech-to-text transcription on the Transcript row.

With transcription enabled, every new video you index has its audio extracted and transcribed automatically. The process runs in two phases, audio extraction followed by transcription, with progress displayed in the indexing window.

Choosing the language

With transcription enabled, the Spoken language list below the checkbox sets the language. It starts on Auto-detect: the speech-to-text model works out the spoken language itself for each video, so a catalog with several languages in it needs no setting changed between them.

Auto-detect decides from the start of each video. A video that opens with silence, music or a tone can be taken for the wrong language, and the whole transcript then comes out as repeated nonsense in that language. When the videos you are about to index all speak the same language, choose it in the list instead. The setting applies to videos you queue after changing it; videos already waiting in the indexing queue keep the language they were queued with.

Videos with several audio tracks

Some videos carry more than one audio track. Which one is transcribed depends on what the tracks are:

  • Tracks in different languages, such as a film with an English and a German track. When you have chosen a Spoken language, the track tagged with that language is transcribed. On Auto-detect the first track is used, as there is nothing to choose by.
  • One recording spread over several tracks, which is how many cameras and external recorders record: MXF files from Sony cameras hold eight separate mono tracks, with the microphone on one of the first four, and MOV and MP4 files can be laid out the same way. The first track is used unless it holds no sound, in which case the following tracks are tried and the first one with sound is transcribed. This takes a second pass over the file, so such videos take longer to transcribe.

What decides which languages are possible is the model, and this is the setting most people need to change first.

English only, or all languages

Fast Video Cataloger ships with the Speech-to-text Base - English only model. It is the smaller and faster of the two, and it can only produce English. Giving it a video in another language does not fail with an error - it transcribes it as though it were English, and the result is nonsense.

For any other language, install the multilingual model:

  1. On the front screen, choose Manage AI models.
  2. Find Speech-to-text Small - all languages and click Download. It is about 490 MB.
  3. When the download finishes the row shows Active: that is the model transcription uses from then on. To switch back to a model you already have, click Use on its row.

The multilingual model covers around a hundred languages, including Hebrew, Arabic, Hindi, Swedish, Polish and Turkish, and it is more accurate than the bundled model on English as well. It is slower, and larger to download, which is why it is not the one that ships in the box.

While an English-only model is active, Preferences shows a warning under the transcription checkbox to say so, and the Spoken language list is fixed on English. If transcripts are coming out as gibberish or as the wrong language, that warning is the first thing to check.

Transcribe Already-Indexed Videos

You can transcribe videos that are already in your catalog without re-indexing their frames. Right-click one or more selected videos and choose Index > Transcribe Audio from the context menu. This runs speech-to-text on the selected videos using your current transcription settings while skipping frame capture entirely.

Transcribe From a Separate Audio File

When the sound was recorded apart from the picture - a separate recorder, a lavalier, an intercom or headset recording - the video's own audio can be too noisy to transcribe. Transcribe the separate recording instead and the transcript is stored on the video.

Select the video. If it has no transcript, the transcript window shows an Add button. Click it and choose the audio file instead of a subtitle file; WAV, MP3, FLAC and OGG files are transcribed. The transcription is queued and runs like Transcribe Audio, with the model and Spoken language from Preferences. To replace an existing transcript, remove it first with Remove transcript.

The transcript times are taken from the audio file. If the recorder was started before or after the camera, the lines are shifted by that difference.

Scripts can do the same with IVideoIndexer.TranscribeFromAudio, for example to transcribe every video in a folder from the WAV file of the same name beside it.

Import SRT Files

There are two ways to import subtitle files as transcripts.

Automatic import during indexing -- In Preferences > Video Indexer, check Add transcript from srt file. When indexing a video, the program looks for a .srt file with the same name as the video in the same folder and imports it as the transcript automatically.

Manual import from the transcript window -- If a video has no transcript, the transcript window shows an Add button. Click it to browse for a subtitle file. Both .srt and .sub formats are supported. Choosing an audio file instead transcribes it, see above. Importing a file replaces any existing transcript for that video.

Using the Transcript Window

The transcript window displays the full transcript for the currently selected video as a scrollable list of lines with timestamps.

Click to seek -- Click any line in the transcript to jump the video player to that point in the video.

Live highlighting -- During playback, the currently spoken line is automatically selected and scrolled into view, so the transcript follows along as the video plays.

Thumbnail tooltips -- Hover over a transcript line to see a tooltip with a thumbnail image from that point in the video, the subtitle text, and the time range.

In-window search -- Use the search bar at the top of the transcript window to find words in the transcript. Type a term and press Enter or click Search. Matching lines are highlighted with a blue background and underline. When there are multiple matches, use the < and > buttons to navigate between them. Each navigation step also seeks the video to the matching line.

Context menu on a line -- Right-click a transcript line to access Add frame (captures a thumbnail at that time) and Add to playlist (adds that time to the current playlist).

Context menu on empty area -- Right-click the empty area of the transcript list to access Remove transcript, which deletes the entire transcript for the video.

Searching Transcripts

The Search window includes a Transcript filter section. Type a word to find all videos that contain that word in their transcript. The transcript search matches the whole word, not partial matches.

When you select a video from the search results, the transcript window opens with the search term pre-filled and matching lines highlighted, so you can immediately see where the word appears and navigate to those moments.


For the other indexing settings, see the Video Indexer preferences page.