Skip to content
All media models/WaveSpeedAI/wavespeed-ai/zonos2
wavespeed / audio

wavespeed-ai/zonos2

Zonos2 is a fast multilingual voice-cloning text-to-speech model that generates natural speech from text using a short reference audio sample. Ready-to-use REST inference API for voice cloning, multilingual TTS, narration, dubbing, character dialogue, virtual assistants, creator content, and professional speech generation workflows with simple integration, no coldstarts, and affordable pricing.

wavespeed-ai/zonos2

Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.

Pricing, with its conditions.

These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.

Source rateBilling unitConditionsEvidence
0.0100 USDsource-listed starting priceDisplayed catalog starting price only. Final charge depends on duration, resolution, outputs and references. Not a fixed per-second or per-image quote.Public structured source

Supported settings from the source

categoryaudio-to-audio

Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.