Skip to content
All media models/WaveSpeedAI/chatterbox/speech-to-speech
wavespeed / audio

chatterbox/speech-to-speech

Chatterbox Speech to Speech is a fast AI voice conversion model that converts source audio into a target voice style with optional reference audio guidance. Ready-to-use REST inference API for voice conversion, speech style transfer, dubbing, character voices, creator content, audio localization, and professional speech-to-speech workflows with simple integration, no coldstarts, and affordable pricing.

chatterbox/speech-to-speech

Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.

Pricing, with its conditions.

These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.

Source rateBilling unitConditionsEvidence
0.0200 USDsource-listed starting priceDisplayed catalog starting price only. Final charge depends on duration, resolution, outputs and references. Not a fixed per-second or per-image quote.Public structured source

Supported settings from the source

categoryaudio-to-audio

Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.