AIAXIO-AI Matched To Your Need

15,503 AI tools for 3,274 Tasks

ElevenLabs Scribe logo

ElevenLabs Scribe

2

17

0

Transcription
The most precise speech-to-text models available.
Input:
Output:
ElevenLabs Scribe screenshot
Updated: Jan 9, 2026 Free + from $5/month

Description

ElevenLabs Speech to Text excels at transforming spoken words into written text with a high degree of precision across a variety of situations and languages.

It offers two primary functionalities: Scribe v2 and Scribe v2 Realtime. The former is designed for converting audio and video into text, making it suitable for generating captions, subtitles, and editable transcripts for various types of recorded media.

It is notable for its capability to accurately transcribe specific words based on context, highlight sound occurrences in transcripts, and identify and label each participant in a conversation.

The latter, Scribe v2 Realtime, is tailored for real-time uses such as live calls, meetings, or AI systems needing immediate transcription.

It employs a streaming-focused design to deliver real-time results while maintaining accuracy. It also incorporates features like accurate speech segmentation for smoother live processing and voice activity detection.

Both Scribe versions are compatible with more than 90 languages and can be integrated into your products using its API.

Pricing Plans

Model
freemium
Packages
1 Package
Price Start From
$5/month
Payment Model
Not specified

Releases

Live real-time speech-to-text model - crafted for streaming transcription with extremely low latency (~150 ms) for live voice interactions.

Ultra-low latency performance - instant speech transcription perfect for conversational AI, voice agents, meetings, and live captioning. 

High precision across numerous languages - compatible with 90+ languages with strong practical performance and benchmark scores. 

Predictive streaming (“negative latency”) - anticipates upcoming words and punctuation to minimize delays. 

Automatic language detection - the model identifies and switches languages during a conversation. 

Advanced streaming controls - including manual commit control, text conditioning, and voice activity detection (VAD). 

Broad audio format compatibility - compatible with PCM (8–48 kHz) and μ-law audio for versatility across various uses.

Reviews

Pros & Cons

Pros

Multilingual transcription

Real-time transcription

Supports 90+ languages

Cons

No offline support

Doesn't support all languages

No free tier

Q&A

New Released

New Released