replicate / audio
realtime-tts-2
Most expressive text-to-speech model from Inworld, with natural-language steering, real-time latency, and multilingual support across 100+ languages.
inworld/realtime-tts-2
Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.
Pricing, with its conditions.
These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.
| Source rate | Billing unit | Conditions | Evidence |
|---|---|---|---|
| $0.025 USD | input character | or 40,000 characters for $1 | Public structured source |
Supported settings from the source
| language | Language of the input text. Use 'auto' to let the model detect the language. Supported production languages: English (en), Chinese (zh), Japanese (ja), Korean (ko), Russian (ru), Italian (it), Spanish (es), Portuguese (pt), French (fr), German (de), Polish (pl), Dutch (nl), Hindi (hi), Hebrew (he), Arabic (ar). · Options: auto, en, zh, ja, ko, ru, it, es, pt, fr, de, pl, nl, hi, he, ar · Default: auto |
|---|---|
| voice_id | The voice to use. Use a preset voice name (e.g. 'Ashley', 'Dennis', 'Alex', 'Darlene') or a custom cloned voice ID. · Default: Ashley |
| sample_rate | Audio sample rate in Hz. · Options: 8000, 16000, 22050, 24000, 32000, 44100, 48000 · Default: 48000 |
| audio_format | Output audio format. · Options: mp3, wav, ogg_opus, flac · Default: mp3 |
Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.