Skip to content
All media models/WaveSpeedAI/microsoft/vibevoice
wavespeed / audio

microsoft/vibevoice

Microsoft VibeVoice text-to-speech model generates long-form speech from text with multi-speaker dialogue support. Choose from 9 voice presets across English, Chinese, and Hindi. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

microsoft/vibevoice

Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.

Pricing, with its conditions.

These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.

Source rateBilling unitConditionsEvidence
0.1200 USDsource-listed starting priceDisplayed catalog starting price only. Final charge depends on duration, resolution, outputs and references. Not a fixed per-second or per-image quote.Public structured source

Supported settings from the source

categorytext-to-audio

Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.