MiniMax H3 Text to Video generates coherent 2K videos from text prompts, with flexible 5-15 second duration and adaptive or custom aspect ratios for cinematic scenes, creative videos, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
minimax/h3/text-to-video
Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.
Inside a quiet mountain tea house, steam curls from a ceramic cup on a rain-speckled window ledge. Focus slowly shifts from the cup to distant fog moving between pine trees. Warm interior light against cool blue rain, restrained live-action cinematography, locked camera, believable steam. No people or lettering.
A red paper lantern hangs from a timber eave during gentle snowfall at dusk. It sways slightly in the wind while warm light softly glows through damp paper. Fixed close camera, snowflakes cross different focal planes, distant rooftops dissolve into blue haze. Natural motion, no symbols or text.
A curled young fern frond slowly unfurls in a damp forest close-up. Tiny dew droplets sparkle along the edge, dark moss fills the background. Gentle time-lapse botanical motion, stable macro camera, realistic green textures and soft morning light. No insects or text.
These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.
Source rate
Billing unit
Conditions
Evidence
0.7000 USD
source-listed starting price
Displayed catalog starting price only. Final charge depends on duration, resolution, outputs and references. Not a fixed per-second or per-image quote.
Public structured source
Supported settings from the source
category
text-to-video
Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.