speech-to-text - Audio Transcription with Whisper Models
Transcribe audio to text using Whisper models via inference.sh CLI
Tags
Updated: 2026-02-23Capabilities
Typical Inputs
Typical Outputs
What this skill does
- transcribe audio
- translate audio
- generate timestamps
- detect language
- extract audio from video
Inputs
- audio files
- audio URLs
- video files
- API key
Outputs
- text transcripts
- JSON transcription data
- timestamped segments
- language detection results
Requirements
- inference.sh CLI installation
- inference.sh login authentication
- valid audio/video input
