THE MODEL CATALOG

Video model APIs & pricing

Find text-to-video, image-to-video and video processing APIs by task. Compare Wan versions, inspect each model’s parameters and submit asynchronous jobs with one Jevrouter key.

Start with a video API guide

Compare version-specific variants, dated pricing examples and working integration code.

532 models · Page 1 of 12

alibaba

happyhorse 1.0 · image to video

Alibaba Happy Horse 1.0 (Image-to-Video) animates a reference image into a cinematic 720p / 1080p video, optionally guided by a text prompt. Smooth camera movement and expressive, stable motion.

Video
Per video · quoted before submission
alibaba

happyhorse 1.0 · reference to video

Alibaba Happy Horse 1.0 (Reference-to-Video) generates new video scenes guided by reference images, maintaining consistent characters, styles, and visual identity. Smooth camera movement and expressive, stable motion.

Video
Per video · quoted before submission
alibaba

happyhorse 1.0 · text to video

Alibaba Happy Horse 1.0 (Text-to-Video) generates cinematic 720p / 1080p videos from text prompts with smooth camera movement, expressive motion, and strong prompt fidelity.

Video
Per video · quoted before submission
alibaba

happyhorse 1.0 · video edit

Alibaba Happy Horse 1.0 (Video Edit) performs prompt-driven video editing with multi-image reference support, supporting 720p/1080p output. Smooth, context-aware edits while preserving motion and temporal consistency.

Video
Per video · quoted before submission
alibaba

happyhorse 1.0 · video extend

Alibaba Happy Horse 1.0 (Video Extend) extends existing videos with seamless AI-generated continuation, supporting 720p/1080p output. Natural, motion-consistent extension.

Video
Per video · quoted before submission
alibaba

happyhorse 1.1 · image to video

Alibaba Happy Horse 1.1 (Image-to-Video) animates a reference image into a cinematic 720p / 1080p video, optionally guided by a text prompt. Smooth camera movement and expressive, stable motion.

Video
Per video · quoted before submission
alibaba

happyhorse 1.1 · reference to video

Alibaba Happy Horse 1.1 (Reference-to-Video) generates new video scenes guided by reference images, maintaining consistent characters, styles, and visual identity. Smooth camera movement and expressive, stable motion.

Video
Per video · quoted before submission
alibaba

happyhorse 1.1 · text to video

Alibaba Happy Horse 1.1 (Text-to-Video) generates cinematic 720p / 1080p videos from text prompts with smooth camera movement, expressive motion, and strong prompt fidelity.

Video
Per video · quoted before submission
alibaba

happyhorse 1.1 · video extend

Alibaba Happy Horse 1.1 (Video Extend) extends existing videos with seamless AI-generated continuation, supporting 720p/1080p output. Natural, motion-consistent extension.

Video
Per video · quoted before submission
alibaba

Wan 2.5 · image to video fast

Alibaba WAN 2.5 Fast converts text or images into synchronized-audio videos in 480p, 720p, or 1080p, offering faster, more affordable generation compared to Google Veo3. REST API, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.5 · image to video

Alibaba WAN 2.5 converts text or images into videos (480p/720p/1080p) with synced audio, faster and more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.5 · text to video fast

Alibaba WAN 2.5 Fast creates synchronized-audio videos from text or images in 720p, faster and more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.5 · text to video

Alibaba WAN 2.5 makes 480p-1080p text/image-to-video with synced audio and is faster, more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.5 · video extend fast

Alibaba WAN 2.5 Fast is an AI-powered video extender that turns short clips into longer videos while preserving audio tracks for seamless extensions. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.5 · video extend

Alibaba WAN 2.5 Video-Extend turns short clips into longer videos with preserved or generated synchronized audio for continuity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.6 · image to video Flash

Alibaba WAN 2.6 Flash converts images into videos (720p/1080p) with optional audio, optimized for speed and cost. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.6 · image to video Pro

Alibaba WAN 2.6 Pro converts images into ultra-high-resolution videos (1080p/2K/4K) with cinematic detail and smooth motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.6 · image to video spicy

Alibaba WAN 2.6 Spicy converts images into unlimited high-quality videos with smooth animations optimized for scalable content generation. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.6 · image to video

Alibaba WAN 2.6 converts text or images into videos (720p/1080p) with synced audio, faster and more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.6 · reference to video Flash

Alibaba WAN 2.6 Reference-to-Video Flash turns character, prop, or scene references from images or videos into new video shots with preserved identity, style, and layout plus smooth, coherent motion. Flash version with faster generation speed. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.6 · reference to video

Alibaba WAN 2.6 Reference-to-Video turns character, prop, or scene references—single or multi-view—into new video shots with preserved identity, style, and layout plus smooth, coherent motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.6 Text-to-Video

Alibaba WAN 2.6 Text-to-Video turns plain prompts into coherent, cinematic clips with crisp detail, stable motion, and strong instruction-following—great for ads, explainers, and social posts. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.6 · video extend

Alibaba WAN 2.6 Video-Extend turns short clips into longer videos with preserved or generated synchronized audio for continuity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.7 · image to video Pro

Alibaba WAN 2.7 Pro converts images into ultra-high-resolution videos (1080p/2K/4K) with cinematic detail and smooth motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.7 · image to video spicy

Alibaba WAN 2.7 Spicy converts images into unlimited high-quality videos with smooth animations optimized for scalable content generation. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.7 · image to video

Alibaba WAN 2.7 converts images into videos (720p/1080p) with optional audio, supporting first and last frame control. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.7 · reference to video

Alibaba WAN 2.7 Reference-to-Video turns character, prop, or scene references from images or videos into new video shots with preserved identity, style, and layout plus smooth, coherent motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.7 · text to video

Alibaba WAN 2.7 Text-to-Video turns plain prompts into coherent, cinematic clips with crisp detail, stable motion, and strong instruction-following—great for ads, explainers, and social posts. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.7 · video edit

Alibaba WAN 2.7 Video Edit performs prompt-driven video editing with multi-image reference support, supporting 720p/1080p output. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 2.7 · video extend

Alibaba WAN 2.7 Video Extend extends existing videos with optional last frame control and audio support, supporting 720p/1080p output. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 3.0 · image to video spicy

Alibaba WAN 3.0 Spicy Image-to-Video converts a first-frame image into a video with optional last-frame guidance, flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 3.0 · image to video

Alibaba WAN 3.0 Image-to-Video converts a first-frame image into a video with optional last-frame guidance, flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 3.0 prime · image to video spicy

Alibaba WAN 3.0 Prime Spicy Image-to-Video converts a first-frame image into a video with optional last-frame guidance, flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 3.0 prime · image to video

Alibaba WAN 3.0 Prime Image-to-Video converts a first-frame image into a video with optional last-frame guidance, flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 3.0 prime · reference to video

Alibaba WAN 3.0 Prime Reference-to-Video combines reference images, videos, and audio with prompts to create coherent videos with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 3.0 prime · text to video

Alibaba WAN 3.0 Prime Text-to-Video generates videos from text prompts with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 3.0 prime · video edit

Wan 3.0 Prime Video-Edit edits an input video using a text prompt and optional reference images and audio. Inputs longer than 15 seconds are trimmed to the first 15 seconds. Output aspect ratio is explicitly selected from the input display dimensions automatically.

Video
Per video · quoted before submission
alibaba

Wan 3.0 prime · video extend

Wan 3.0 Prime Video Extend continues a video from its final frame and appends a newly generated 2-30 second segment. Inputs longer than 120 seconds retain their last 120 seconds. Supports 480p, 720p, and 1080p output.

Video
Per video · quoted before submission
alibaba

Wan 3.0 · reference to video

Alibaba WAN 3.0 Reference-to-Video combines reference images, videos, and audio with prompts to create coherent videos with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 3.0 · text to video

Alibaba WAN 3.0 Text-to-Video generates videos from text prompts with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
alibaba

Wan 3.0 · video edit

Wan 3.0 Video-Edit edits an input video using a text prompt and optional reference images and audio. Inputs longer than 15 seconds are trimmed to the first 15 seconds. Output aspect ratio is explicitly selected from the input display dimensions automatically.

Video
Per video · quoted before submission
alibaba

Wan 3.0 · video extend

Wan 3.0 Video Extend continues a video from its final frame and appends a newly generated 2-30 second segment. Inputs longer than 120 seconds retain their last 120 seconds. Supports 480p, 720p, and 1080p output.

Video
Per video · quoted before submission
black-forest-labs

flux 3 · image to video draft

Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .

Video
Per video · quoted before submission
black-forest-labs

flux 3 · image to video

Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .

Video
Per video · quoted before submission
black-forest-labs

flux 3 · start end to video draft

Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .

Video
Per video · quoted before submission
black-forest-labs

flux 3 · start end to video

Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .

Video
Per video · quoted before submission
black-forest-labs

flux 3 · text to video draft

Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .

Video
Per video · quoted before submission
black-forest-labs

flux 3 · text to video

Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .

Video
Per video · quoted before submission

Choose a model. Keep one video endpoint.

Pin a variant or let Auto compare compatible tasks within your key’s model pool and reservation limit.

Video API tutorial

Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.