Skip to content
All media models/Replicate/speech-02-turbo
replicate / audio

speech-02-turbo

Text-to-Audio (T2A) that offers voice synthesis, emotional expression, and multilingual capabilities. Designed for real-time applications with low latency

minimax/speech-02-turbo

Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.

Pricing, with its conditions.

These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.

Source rateBilling unitConditionsEvidence
$0.06 USDinput tokenor around 16,666 tokens for $1Public structured source

Supported settings from the source

voice_idVoice to synthesize. Pick any MiniMax system voice (e.g. English_Wiselady, English_Deep-VoicedGentleman) or a voice_id returned by https://replicate.com/minimax/voice-cloning. See the full list of voices in the README. · Default: English_Wiselady
sample_rateAudio sample rate in Hz. · Options: 8000, 16000, 22050, 24000, 32000, 44100 · Default: 32000
audio_formatFile format for the generated audio. Choose mp3 for general use, wav/flac for lossless, or pcm for raw bytes. · Options: mp3, wav, flac, pcm · Default: mp3
language_boostOptional language hint. Choose Automatic to let MiniMax detect the language, or pick a specific locale. · Options: None, Automatic, Chinese, Chinese,Yue, Cantonese, English, Arabic, Russian, Spanish, French, Portuguese, German, Turkish, Dutch, Ukrainian, Vietnamese, Indonesian, Japanese, Italian, Korean, Thai, Polish, Romanian, Greek, Czech, Finnish, Hindi, Bulgarian, Danish, Hebrew, Malay, Persian, Slovak, Swedish, Croatian, Filipino, Hungarian, Norwegian, Slovenian, Catalan, Nynorsk, Tamil, Afrikaans · Default: None

Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.

speech-02-turbo · replicate API pricing · Vidily