Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
audio recognition detection speech speech-recognition filler labeling transcription whisper speech-processing asr timestamps verbatim stuttering turn-taking stutter-detection full-duplex-audio
-
Updated
Aug 23, 2026 - Python