Skip to content
replicate / audio

realtime-tts-2

Most expressive text-to-speech model from Inworld, with natural-language steering, real-time latency, and multilingual support across 100+ languages.

inworld/realtime-tts-2

Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.

Pricing, with its conditions.

These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.

Source rateBilling unitConditionsEvidence
$0.025 USDinput characteror 40,000 characters for $1Public structured source

Supported settings from the source

languageLanguage of the input text. Use 'auto' to let the model detect the language. Supported production languages: English (en), Chinese (zh), Japanese (ja), Korean (ko), Russian (ru), Italian (it), Spanish (es), Portuguese (pt), French (fr), German (de), Polish (pl), Dutch (nl), Hindi (hi), Hebrew (he), Arabic (ar). · Options: auto, en, zh, ja, ko, ru, it, es, pt, fr, de, pl, nl, hi, he, ar · Default: auto
voice_idThe voice to use. Use a preset voice name (e.g. 'Ashley', 'Dennis', 'Alex', 'Darlene') or a custom cloned voice ID. · Default: Ashley
sample_rateAudio sample rate in Hz. · Options: 8000, 16000, 22050, 24000, 32000, 44100, 48000 · Default: 48000
audio_formatOutput audio format. · Options: mp3, wav, ogg_opus, flac · Default: mp3

Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.

realtime-tts-2 · replicate API pricing · Vidily