wan 2.2
wan 2.2, image-to-video, 5.0s-480pWan-S2V is a video model that generates high-quality videos from static images and audio, with realistic facial expressions, body movements, and professional camera work for film and television applications
fal-ai/wan/v2.2-14b/speech-to-video
Source checked: 2026-09-13 · Provider endpoint, not an independently benchmarked model.
These are source observations, not interchangeable quotes. Token, compute-time, output-duration and per-request prices cannot be compared by the number alone. Input charges and optional features may add to the total.
Your request will cost $0.20 per video second for 720p, $0.15 per video second for 580p, $0.10 per video second for 480p. Video seconds are calculated at 16 frames per second.
| category | audio-to-video |
|---|---|
| license | commercial |
| audio_url | The URL of the audio file. |
| num_frames | Number of frames to generate. Must be between 40 to 120, (must be multiple of 4). · Min: 40 · Max: 120 · Default: 80 |
| resolution | Resolution of the generated video (480p, 580p, or 720p). · Options: 480p, 580p, 720p · Default: 480p |
| video_quality | The quality of the output video. Higher quality means better visual quality but larger file size. · Options: low, medium, high, maximum · Default: high |
| frames_per_second | Frames per second of the generated video. Must be between 4 to 60. When using interpolation and `adjust_fps_for_interpolation` is set to true (default true,) the final FPS will be multiplied by the number of interpolated frames plus one. For example, if the generated frames per second is 16 and the number of interpolated frames is 1, the final frames per second will be 32. If `adjust_fps_for_interpolation` is set to false, this value will be used as-is. · Default: 16 |
Unlisted settings, region availability, deposits, taxes and commercial terms remain unverified. Confirm the exact endpoint before purchase.
Fast, Pro, editing and audio variants remain separate. Match the exact configuration before comparing costs.
Source-listed endpoints and pricing conditions. Variants are not unique base models; account quotes do not enter cheapest-price rankings.
Counts describe collected endpoints, not unique models or a guarantee of complete coverage. Price evidence may be a numeric rate or source billing text. Providers with no published endpoints are still awaiting collection.
58 endpoints · 1/3
Variants listed separately. Open a model to check units and conditions.
wan 2.2, image-to-video, 5.0s-480pwan 2.2, image-to-video, 5.0s-480p
wan 2.2, image-to-video, 5.0s-720pwan 2.2, image-to-video, 5.0s-720p
wan 2.2, image-to-video, 5.0s-580pwan 2.2, image-to-video, 5.0s-580p
wan 2.2, text-to-video, 5.0s-580pwan 2.2, text-to-video, 5.0s-580p
wan 2.2, text-to-video, 5.0s-480pwan 2.2, text-to-video, 5.0s-480p
wan 2.2, text-to-video, 5.0s-720pwan 2.2, text-to-video, 5.0s-720p
Wan 2.2 A14B Turbo API Speech to Video, 480pWan 2.2 A14B Turbo API Speech to Video, 480p
Wan 2.2 A14B Turbo API Speech to Video, 720pWan 2.2 A14B Turbo API Speech to Video, 720p
Wan 2.2 A14B Turbo API Speech to Video, 580pWan 2.2 A14B Turbo API Speech to Video, 580p
wan 2.2 Animate, 2.2 Animate Replace, 1.0s-720pwan 2.2 Animate, 2.2 Animate Replace, 1.0s-720p
wan 2.2 Animate, 2.2 Animate Replace, 1.0s-580pwan 2.2 Animate, 2.2 Animate Replace, 1.0s-580p
wan 2.2 Animate, 2.2 Animate Replace, 1.0s-480pwan 2.2 Animate, 2.2 Animate Replace, 1.0s-480p
wan 2.2 Animate, 2.2 Animate Move, 1.0s-480pwan 2.2 Animate, 2.2 Animate Move, 1.0s-480p
wan 2.2 Animate, 2.2 Animate Move, 1.0s-580pwan 2.2 Animate, 2.2 Animate Move, 1.0s-580p
wan 2.2 Animate, 2.2 Animate Move, 1.0s-720pwan 2.2 Animate, 2.2 Animate Move, 1.0s-720p
fal-ai/wan-22-vace-fun-a14b/depthVACE Fun for Wan 2.2 A14B from Alibaba-PAI
fal-ai/wan-22-vace-fun-a14b/inpaintingVACE Fun for Wan 2.2 A14B from Alibaba-PAI
fal-ai/wan-22-vace-fun-a14b/outpaintingVACE Fun for Wan 2.2 A14B from Alibaba-PAI
fal-ai/wan-22-vace-fun-a14b/reframeVACE Fun for Wan 2.2 A14B from Alibaba-PAI
fal-ai/wan/v2.2-14b/animate/moveWan-Animate is a video model that generates high-fidelity character videos by replicating the expressions and movements of characters from reference videos.
fal-ai/wan/v2.2-14b/animate/replaceWan-Animate Replace is a model that can integrate animated characters into reference videos, replacing the original character while preserving the scene’s lighting and color tone for seamless environm
fal-ai/wan/v2.2-14b/speech-to-videoWan-S2V is a video model that generates high-quality videos from static images and audio, with realistic facial expressions, body movements, and professional camera work for film and television applic
Model sample · hivamohfal-ai/wan/v2.2-a14b/text-to-videoWan-2.2 text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.
fal-ai/wan/v2.2-a14b/text-to-video/loraWan-2.2 text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts. This endpoint supports LoRAs made for Wan 2.2.