replicate / audio
play-dialog
End-to-end AI speech model designed for natural-sounding conversational speech synthesis, with support for context-aware prosody, intonation, and emotional expression.
playht/play-dialog
Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.
Pricing, with its conditions.
These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.
| Source rate | Billing unit | Conditions | Evidence |
|---|---|---|---|
| $1 USD | second of output audio | or 1,000 seconds for $1 | Public structured source |
Supported settings from the source
| voice | Voice to use for generation · Options: Angelo (Young male US conversational voice), Arsenio (Middle-aged male US African American conversational voice), Cillian (Middle-aged male Irish conversational voice), Timo (Middle-aged male US conversational voice), Dexter (Middle-aged male US conversational voice), Miles (Young male US African American conversational voice), Briggs (Elderly male US Southern (Oklahoma) conversational voice), Deedee (Middle-aged female US African American conversational voice), Nia (Young female US conversational voice), Inara (Middle-aged female US African American conversational voice), Constanza (Young female US Latin American conversational voice), Gideon (Elderly male British narrative voice), Casper (Middle-aged male US narrative voice), Mitch (Middle-aged male Australian narrative voice), Ava (Middle-aged female Australian narrative voice), Carmen (Middle-aged female Spanish conversational voice, calm and warm), Andrei (Middle-aged male Russian conversational voice, calm and warm), Ilias (Middle-aged male German narrative voice, deep and calm), Gaelle (Middle-aged female French conversational voice, professional), Alessandro (Older male Italian conversational voice, warm and gravelly), Yumiko (Young female Japanese narrative voice, warm and light) · Default: Angelo (Young male US conversational voice) |
|---|---|
| voice_2 | Optional second voice to use for generation · Options: None, Angelo (Young male US conversational voice), Arsenio (Middle-aged male US African American conversational voice), Cillian (Middle-aged male Irish conversational voice), Timo (Middle-aged male US conversational voice), Dexter (Middle-aged male US conversational voice), Miles (Young male US African American conversational voice), Briggs (Elderly male US Southern (Oklahoma) conversational voice), Deedee (Middle-aged female US African American conversational voice), Nia (Young female US conversational voice), Inara (Middle-aged female US African American conversational voice), Constanza (Young female US Latin American conversational voice), Gideon (Elderly male British narrative voice), Casper (Middle-aged male US narrative voice), Mitch (Middle-aged male Australian narrative voice), Ava (Middle-aged female Australian narrative voice), Carmen (Middle-aged female Spanish conversational voice, calm and warm), Andrei (Middle-aged male Russian conversational voice, calm and warm), Ilias (Middle-aged male German narrative voice, deep and calm), Gaelle (Middle-aged female French conversational voice, professional), Alessandro (Older male Italian conversational voice, warm and gravelly), Yumiko (Young female Japanese narrative voice, warm and light) · Default: None |
| language | The language of the text to be spoken. · Options: afrikaans, albanian, amharic, arabic, bengali, bulgarian, catalan, croatian, czech, danish, dutch, english, french, galician, german, greek, hebrew, hindi, hungarian, indonesian, italian, japanese, korean, malay, mandarin, polish, portuguese, russian, serbian, spanish, swedish, tagalog, thai, turkish, ukrainian, urdu, xhosa · Default: english |
| voice_conditioning_seconds | The number of seconds of conditioning to use from the selected voice. Lower values generate audio less similar to the cloned voice, but lead to more model stability and expressiveness. Higher values create output more similar to the cloned voice, but can lead to model instability and reduced expressiveness. · Min: 1 · Max: 60 · Default: 20 |
| voice_conditioning_seconds_2 | The number of seconds of conditioning to use from the second selected voice. · Min: 1 · Max: 60 · Default: 20 |
Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.