跳至正文
全部媒体模型/Replicate/grok-text-to-speech
replicate / audio

grok-text-to-speech

Convert text to natural-sounding speech with xAI's Grok TTS. 5 voices, 20 languages, expressive speech tags, and high-fidelity MP3 / WAV / telephony audio output.

xai/grok-text-to-speech

来源检查时间: 2026-09-13 · 供应商端点记录,不代表独立效果评测。

价格与适用条件。

以下是来源观察记录,不能直接相互替代。Token、计算耗时、输出时长与按次价格不能仅比较数值;输入和可选功能可能另行计费。

来源标价计费单位条件依据
$0.015 USDinput characteror around 66,666 characters for $1公开结构化来源

来源提供的参数

voiceVoice to use for synthesis. 'eve' is energetic and upbeat (default), 'ara' is warm and friendly, 'rex' is confident and clear, 'sal' is smooth and balanced, 'leo' is authoritative and strong. · Options: eve, ara, rex, sal, leo · Default: eve
languageBCP-47 language code for the input text. Set to 'auto' to let the model auto-detect the language. · Options: auto, en, ar-EG, ar-SA, ar-AE, bn, zh, fr, de, hi, id, it, ja, ko, pt-BR, pt-PT, ru, es-MX, es-ES, tr, vi · Default: auto
sample_rateAudio sample rate in Hz. Higher rates produce better quality at the cost of file size. · Options: 8000, 16000, 22050, 24000, 44100, 48000 · Default: 24000
output_formatAudio codec. 'mp3' is best for general use, 'wav' for lossless audio, 'pcm' for raw audio pipelines, 'mulaw'/'alaw' for telephony. · Options: mp3, wav, pcm, mulaw, alaw · Default: mp3

未列出的参数、地区可用性、最低充值、税费和商用条款仍待核实,购买前请核对精确端点。

grok-text-to-speech · replicate API 价格 · Vidily