✅ Supports common formats: MP4 / MOV / WEBM / MP3 / WAV
✅ Flow: Extract audio -> Speech recognition -> Output text
ℹ️ Note: Whisper tiny is currently used and has limits for homophones and colloquial speech.
ℹ️ Tip: Click "Start Speech to Text" first, then open "Doubao Web" to paste and refine.
Select video or audio file
Click or drag video / audio
Supports MP4/MOV/WEBM/MP3/WAV | First run loads FFmpeg; model source can switch local/r2
Initializing...