Speech Recognition
Recognize system audio and display the source text or translation as subtitles
Speech Recognition is designed for videos, live streams, meetings, and other audio. It recognizes speech from system audio, can translate the result, and can display subtitles in a floating window.
Workflow
- Download or import an ASR model under
Setting > Speech. - Open this page and select the recognition language and an installed model.
- To translate recognition results, enable
Enable Translationand choose the target language and an AI model or machine translation service. - Start recognition, then open the
Subtitle Windowwhen you want subtitles over other applications.
Recognition options
| Option | Purpose |
|---|---|
| Audio Source | Use general system audio or isolate a specific process to reduce interference. |
| Recognition language | Match the language spoken in the audio for better recognition accuracy. |
| ASR model | Select an installed recognition model compatible with the language. |
| Enable Translation and Target Lang | Control whether translated subtitles are generated and which language they use. |
| Translation engine and model | Select an AI model or configured machine translation service. |
| Prompt | Select a prompt for AI-based speech translation. |
| Real-time Preview | Translate partial recognition results. Disable it to translate confirmed sentences only. |
Subtitle window
The page also configures subtitle appearance presets, primary and secondary subtitle colors and sizes, background, display mode, automatic clearing, sentences per line, and completed-line history. The floating window can be locked for click-through use and its history can be cleared manually.
This feature is available on Windows only.