LogoClawIndex
CasesSkillsAbout
LogoClawIndex

whisper - Robust Multilingual Speech Recognition and Translation

Transcribes audio into text and translates speech into English across 99 languages using the Whisper model.

Tags

Updated: 2026-09-18
WhisperSpeech RecognitionASRSpeech-To-TextTranscriptionTranslationAudio Processing

Capabilities

Transcribe audio into textTranslate speech into English textGenerate word level timestampsExport transcriptions to subtitle files

Typical Inputs

Audio filesVideo filesLanguage codes

Typical Outputs

Transcribed textTXT text filesSRT subtitle files

What this skill does

  • Transcribe audio into text
  • Translate speech into English text
  • Generate word level timestamps
  • Export transcriptions to subtitle files
  • Execute transcription via command line

Inputs

  • Audio files
  • Video files
  • Language codes
  • Initial prompt text

Outputs

  • Transcribed text
  • TXT text files
  • SRT subtitle files
  • WebVTT subtitle files
  • JSON output files

Requirements

  • Python 3.8-3.11
  • ffmpeg
  • openai-whisper package
  • Linux or macOS platform

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.