The speech model is not ready
Open Settings and let Desktop finish downloading and loading the local model before retrying.
Create TXT transcripts and SRT subtitles from downloaded video on your own computer. Desktop prepares audio with FFmpeg and runs the speech model without uploading the media.
Desktop prepares 16 kHz audio with FFmpeg, detects speech, then recognizes one segment at a time with the local model.
Each step maps to a control or task state in the Desktop application shown above.
Enable TXT + SRT before starting a compatible video download.
On first use, confirm setup and wait until Settings reports Enabled or Ready.
Desktop downloads the video, prepares audio and processes transcription serially.
Find the matching .txt and .srt files beside the downloaded media.
Failures remain attached to their task so completed items in the same queue stay available.
Open Settings and let Desktop finish downloading and loading the local model before retrying.
Confirm the video contains audible speech. Music-only, silent or extremely noisy audio can produce empty files.
Keep the downloaded video available and retry. A damaged or unsupported file may prevent FFmpeg from preparing audio.
Choose the current package for your computer. Release cards appear only when that package is available.
Free: 10 successful single-video parses per day; no profile extraction. Desktop Pro: $9.90/month or $79.90/year for unlimited parsing and supported profile extraction. Compare plans.
No. Audio preparation and speech recognition run locally on the computer.
The current Settings screen lists Mandarin, Cantonese, English, Japanese and Korean.
No. Results depend on language, speech clarity, background noise, overlap and the source recording.