LogoClawIndex
CasesSkillsAbout
LogoClawIndex

speech-to-text - Audio Transcription with Whisper Models

Transcribe audio to text using Whisper models via inference.sh CLI

Tags

Updated: 2026-02-23

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • transcribe audio
  • translate audio
  • generate timestamps
  • detect language
  • extract audio from video

Inputs

  • audio files
  • audio URLs
  • video files
  • API key

Outputs

  • text transcripts
  • JSON transcription data
  • timestamped segments
  • language detection results

Requirements

  • inference.sh CLI installation
  • inference.sh login authentication
  • valid audio/video input

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.
audio
transcription
whisper
speech
translation
subtitles
transcribe audio
translate audio
generate timestamps
detect language
audio files
audio URLs
video files
text transcripts
JSON transcription data
timestamped segments