MODELS & PRICING
The right model.
For your audio.
Explore each deployment’s supported languages, settings and API examples. Prices are in USD per audio minute, read from the same catalog as Console. Models may have different rates; current prices are not a promise of one permanent rate.
Whisper Large V3 Turbo
Fast multilingual speech recognition tuned for strong accuracy and lower latency.
whisper-large-v3-turbo-1Whisper Large V3
Multilingual speech recognition for accurate transcription, translation and timestamps.
whisper-large-v3-1Whisper Large V3 Turbo
Fast multilingual speech recognition tuned for strong accuracy and lower latency.
whisper-large-v3-turbo-2Whisper Large V3
Multilingual speech recognition for accurate transcription, translation and timestamps.
whisper-large-v3-2Whisper Large V3
Multilingual speech recognition for accurate transcription, translation and timestamps.
whisper-large-v3-3Parakeet TDT 0.6B v3
Efficient multilingual ASR designed for high-throughput, low-latency transcription.
parakeet-tdt-0.6b-v3Whisper Realtime
Streaming multilingual speech recognition powered by the Whisper model family.
whisper-realtimeNova-3 Batch
General-purpose speech recognition with language detection and detailed timestamps.
nova-3-batchNova-3 Realtime
Realtime speech recognition with fast partials and production-ready final transcripts.
nova-3-realtimeUniversal-3 Pro Batch
High-accuracy speech recognition for multilingual audio and long-form transcription.
universal-3-pro-batchUniversal-3 Realtime Pro
Low-latency streaming ASR for live conversations and responsive voice applications.
universal-3-realtime-proMOSS Transcribe 1.0
General speech transcription optimized for clear, structured text output.
moss-transcribe-1.0MOSS Transcribe Diarize Pro
Speech transcription with speaker-aware segments for conversations and meetings.
moss-transcribe-diarize-proFun-ASR Flash
Fast multilingual speech recognition for short audio and interactive workloads.
fun-asr-flash-sgQwen Audio 3.0 ASR Flash
Multilingual audio understanding and fast speech-to-text transcription.
qwen-audio-3-asr-flash-sgQwen Audio 3.0 ASR Flash Streaming
Realtime multilingual ASR with incremental transcripts for streaming audio.
qwen-audio-3-asr-streaming-sgQwen3 Forced Alignment
Aligns an existing transcript to speech and returns precise word-level timestamps.
qwenasr-alignDeepFilterNet3
Neural speech enhancement that suppresses noise while preserving voice clarity.
deepfilternet3Silero VAD
Detects speech regions in audio for trimming, segmentation and preprocessing.
silero-vad