whisper - Robust Multilingual Speech Recognition and Translation
Transcribes audio into text and translates speech into English across 99 languages using the Whisper model.
Tags
Updated: 2026-09-18Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Transcribe audio into text
- Translate speech into English text
- Generate word level timestamps
- Export transcriptions to subtitle files
- Execute transcription via command line
Inputs
- Audio files
- Video files
- Language codes
- Initial prompt text
Outputs
- Transcribed text
- TXT text files
- SRT subtitle files
- WebVTT subtitle files
- JSON output files
Requirements
- Python 3.8-3.11
- ffmpeg
- openai-whisper package
- Linux or macOS platform
