Skip to content
replicate / audio

play-dialog

End-to-end AI speech model designed for natural-sounding conversational speech synthesis, with support for context-aware prosody, intonation, and emotional expression.

playht/play-dialog

Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.

Pricing, with its conditions.

These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.

Source rateBilling unitConditionsEvidence
$1 USDsecond of output audioor 1,000 seconds for $1Public structured source

Supported settings from the source

voiceVoice to use for generation · Options: Angelo (Young male US conversational voice), Arsenio (Middle-aged male US African American conversational voice), Cillian (Middle-aged male Irish conversational voice), Timo (Middle-aged male US conversational voice), Dexter (Middle-aged male US conversational voice), Miles (Young male US African American conversational voice), Briggs (Elderly male US Southern (Oklahoma) conversational voice), Deedee (Middle-aged female US African American conversational voice), Nia (Young female US conversational voice), Inara (Middle-aged female US African American conversational voice), Constanza (Young female US Latin American conversational voice), Gideon (Elderly male British narrative voice), Casper (Middle-aged male US narrative voice), Mitch (Middle-aged male Australian narrative voice), Ava (Middle-aged female Australian narrative voice), Carmen (Middle-aged female Spanish conversational voice, calm and warm), Andrei (Middle-aged male Russian conversational voice, calm and warm), Ilias (Middle-aged male German narrative voice, deep and calm), Gaelle (Middle-aged female French conversational voice, professional), Alessandro (Older male Italian conversational voice, warm and gravelly), Yumiko (Young female Japanese narrative voice, warm and light) · Default: Angelo (Young male US conversational voice)
voice_2Optional second voice to use for generation · Options: None, Angelo (Young male US conversational voice), Arsenio (Middle-aged male US African American conversational voice), Cillian (Middle-aged male Irish conversational voice), Timo (Middle-aged male US conversational voice), Dexter (Middle-aged male US conversational voice), Miles (Young male US African American conversational voice), Briggs (Elderly male US Southern (Oklahoma) conversational voice), Deedee (Middle-aged female US African American conversational voice), Nia (Young female US conversational voice), Inara (Middle-aged female US African American conversational voice), Constanza (Young female US Latin American conversational voice), Gideon (Elderly male British narrative voice), Casper (Middle-aged male US narrative voice), Mitch (Middle-aged male Australian narrative voice), Ava (Middle-aged female Australian narrative voice), Carmen (Middle-aged female Spanish conversational voice, calm and warm), Andrei (Middle-aged male Russian conversational voice, calm and warm), Ilias (Middle-aged male German narrative voice, deep and calm), Gaelle (Middle-aged female French conversational voice, professional), Alessandro (Older male Italian conversational voice, warm and gravelly), Yumiko (Young female Japanese narrative voice, warm and light) · Default: None
languageThe language of the text to be spoken. · Options: afrikaans, albanian, amharic, arabic, bengali, bulgarian, catalan, croatian, czech, danish, dutch, english, french, galician, german, greek, hebrew, hindi, hungarian, indonesian, italian, japanese, korean, malay, mandarin, polish, portuguese, russian, serbian, spanish, swedish, tagalog, thai, turkish, ukrainian, urdu, xhosa · Default: english
voice_conditioning_secondsThe number of seconds of conditioning to use from the selected voice. Lower values generate audio less similar to the cloned voice, but lead to more model stability and expressiveness. Higher values create output more similar to the cloned voice, but can lead to model instability and reduced expressiveness. · Min: 1 · Max: 60 · Default: 20
voice_conditioning_seconds_2The number of seconds of conditioning to use from the second selected voice. · Min: 1 · Max: 60 · Default: 20

Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.