Skip to content
All media models/Replicate/grok-text-to-speech
replicate / audio

grok-text-to-speech

Convert text to natural-sounding speech with xAI's Grok TTS. 5 voices, 20 languages, expressive speech tags, and high-fidelity MP3 / WAV / telephony audio output.

xai/grok-text-to-speech

Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.

Pricing, with its conditions.

These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.

Source rateBilling unitConditionsEvidence
$0.015 USDinput characteror around 66,666 characters for $1Public structured source

Supported settings from the source

voiceVoice to use for synthesis. 'eve' is energetic and upbeat (default), 'ara' is warm and friendly, 'rex' is confident and clear, 'sal' is smooth and balanced, 'leo' is authoritative and strong. · Options: eve, ara, rex, sal, leo · Default: eve
languageBCP-47 language code for the input text. Set to 'auto' to let the model auto-detect the language. · Options: auto, en, ar-EG, ar-SA, ar-AE, bn, zh, fr, de, hi, id, it, ja, ko, pt-BR, pt-PT, ru, es-MX, es-ES, tr, vi · Default: auto
sample_rateAudio sample rate in Hz. Higher rates produce better quality at the cost of file size. · Options: 8000, 16000, 22050, 24000, 44100, 48000 · Default: 24000
output_formatAudio codec. 'mp3' is best for general use, 'wav' for lossless audio, 'pcm' for raw audio pipelines, 'mulaw'/'alaw' for telephony. · Options: mp3, wav, pcm, mulaw, alaw · Default: mp3

Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.

grok-text-to-speech · replicate API pricing · Vidily