THE MODEL CATALOG
Video model APIs & pricing
Find text-to-video, image-to-video and video processing APIs by task. Compare Wan versions, inspect each model’s parameters and submit asynchronous jobs with one Jevrouter key.
Start with a video API guide
Compare version-specific variants, dated pricing examples and working integration code.
532 models · Page 1 of 12
happyhorse 1.0 · image to video
Alibaba Happy Horse 1.0 (Image-to-Video) animates a reference image into a cinematic 720p / 1080p video, optionally guided by a text prompt. Smooth camera movement and expressive, stable motion.
happyhorse 1.0 · reference to video
Alibaba Happy Horse 1.0 (Reference-to-Video) generates new video scenes guided by reference images, maintaining consistent characters, styles, and visual identity. Smooth camera movement and expressive, stable motion.
happyhorse 1.0 · text to video
Alibaba Happy Horse 1.0 (Text-to-Video) generates cinematic 720p / 1080p videos from text prompts with smooth camera movement, expressive motion, and strong prompt fidelity.
happyhorse 1.0 · video edit
Alibaba Happy Horse 1.0 (Video Edit) performs prompt-driven video editing with multi-image reference support, supporting 720p/1080p output. Smooth, context-aware edits while preserving motion and temporal consistency.
happyhorse 1.0 · video extend
Alibaba Happy Horse 1.0 (Video Extend) extends existing videos with seamless AI-generated continuation, supporting 720p/1080p output. Natural, motion-consistent extension.
happyhorse 1.1 · image to video
Alibaba Happy Horse 1.1 (Image-to-Video) animates a reference image into a cinematic 720p / 1080p video, optionally guided by a text prompt. Smooth camera movement and expressive, stable motion.
happyhorse 1.1 · reference to video
Alibaba Happy Horse 1.1 (Reference-to-Video) generates new video scenes guided by reference images, maintaining consistent characters, styles, and visual identity. Smooth camera movement and expressive, stable motion.
happyhorse 1.1 · text to video
Alibaba Happy Horse 1.1 (Text-to-Video) generates cinematic 720p / 1080p videos from text prompts with smooth camera movement, expressive motion, and strong prompt fidelity.
happyhorse 1.1 · video extend
Alibaba Happy Horse 1.1 (Video Extend) extends existing videos with seamless AI-generated continuation, supporting 720p/1080p output. Natural, motion-consistent extension.
Wan 2.5 · image to video fast
Alibaba WAN 2.5 Fast converts text or images into synchronized-audio videos in 480p, 720p, or 1080p, offering faster, more affordable generation compared to Google Veo3. REST API, no coldstarts, affordable pricing.
Wan 2.5 · image to video
Alibaba WAN 2.5 converts text or images into videos (480p/720p/1080p) with synced audio, faster and more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Wan 2.5 · text to video fast
Alibaba WAN 2.5 Fast creates synchronized-audio videos from text or images in 720p, faster and more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Wan 2.5 · text to video
Alibaba WAN 2.5 makes 480p-1080p text/image-to-video with synced audio and is faster, more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Wan 2.5 · video extend fast
Alibaba WAN 2.5 Fast is an AI-powered video extender that turns short clips into longer videos while preserving audio tracks for seamless extensions. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Wan 2.5 · video extend
Alibaba WAN 2.5 Video-Extend turns short clips into longer videos with preserved or generated synchronized audio for continuity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Wan 2.6 · image to video Flash
Alibaba WAN 2.6 Flash converts images into videos (720p/1080p) with optional audio, optimized for speed and cost. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.6 · image to video Pro
Alibaba WAN 2.6 Pro converts images into ultra-high-resolution videos (1080p/2K/4K) with cinematic detail and smooth motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.6 · image to video spicy
Alibaba WAN 2.6 Spicy converts images into unlimited high-quality videos with smooth animations optimized for scalable content generation. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.6 · image to video
Alibaba WAN 2.6 converts text or images into videos (720p/1080p) with synced audio, faster and more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.6 · reference to video Flash
Alibaba WAN 2.6 Reference-to-Video Flash turns character, prop, or scene references from images or videos into new video shots with preserved identity, style, and layout plus smooth, coherent motion. Flash version with faster generation speed. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.6 · reference to video
Alibaba WAN 2.6 Reference-to-Video turns character, prop, or scene references—single or multi-view—into new video shots with preserved identity, style, and layout plus smooth, coherent motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.6 Text-to-Video
Alibaba WAN 2.6 Text-to-Video turns plain prompts into coherent, cinematic clips with crisp detail, stable motion, and strong instruction-following—great for ads, explainers, and social posts. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.6 · video extend
Alibaba WAN 2.6 Video-Extend turns short clips into longer videos with preserved or generated synchronized audio for continuity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Wan 2.7 · image to video Pro
Alibaba WAN 2.7 Pro converts images into ultra-high-resolution videos (1080p/2K/4K) with cinematic detail and smooth motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.7 · image to video spicy
Alibaba WAN 2.7 Spicy converts images into unlimited high-quality videos with smooth animations optimized for scalable content generation. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.7 · image to video
Alibaba WAN 2.7 converts images into videos (720p/1080p) with optional audio, supporting first and last frame control. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.7 · reference to video
Alibaba WAN 2.7 Reference-to-Video turns character, prop, or scene references from images or videos into new video shots with preserved identity, style, and layout plus smooth, coherent motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.7 · text to video
Alibaba WAN 2.7 Text-to-Video turns plain prompts into coherent, cinematic clips with crisp detail, stable motion, and strong instruction-following—great for ads, explainers, and social posts. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.7 · video edit
Alibaba WAN 2.7 Video Edit performs prompt-driven video editing with multi-image reference support, supporting 720p/1080p output. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 2.7 · video extend
Alibaba WAN 2.7 Video Extend extends existing videos with optional last frame control and audio support, supporting 720p/1080p output. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 3.0 · image to video spicy
Alibaba WAN 3.0 Spicy Image-to-Video converts a first-frame image into a video with optional last-frame guidance, flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 3.0 · image to video
Alibaba WAN 3.0 Image-to-Video converts a first-frame image into a video with optional last-frame guidance, flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 3.0 prime · image to video spicy
Alibaba WAN 3.0 Prime Spicy Image-to-Video converts a first-frame image into a video with optional last-frame guidance, flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 3.0 prime · image to video
Alibaba WAN 3.0 Prime Image-to-Video converts a first-frame image into a video with optional last-frame guidance, flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 3.0 prime · reference to video
Alibaba WAN 3.0 Prime Reference-to-Video combines reference images, videos, and audio with prompts to create coherent videos with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 3.0 prime · text to video
Alibaba WAN 3.0 Prime Text-to-Video generates videos from text prompts with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 3.0 prime · video edit
Wan 3.0 Prime Video-Edit edits an input video using a text prompt and optional reference images and audio. Inputs longer than 15 seconds are trimmed to the first 15 seconds. Output aspect ratio is explicitly selected from the input display dimensions automatically.
Wan 3.0 prime · video extend
Wan 3.0 Prime Video Extend continues a video from its final frame and appends a newly generated 2-30 second segment. Inputs longer than 120 seconds retain their last 120 seconds. Supports 480p, 720p, and 1080p output.
Wan 3.0 · reference to video
Alibaba WAN 3.0 Reference-to-Video combines reference images, videos, and audio with prompts to create coherent videos with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 3.0 · text to video
Alibaba WAN 3.0 Text-to-Video generates videos from text prompts with flexible 2-30 second duration, resolution, aspect ratio, audio, and deep-thinking controls. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Wan 3.0 · video edit
Wan 3.0 Video-Edit edits an input video using a text prompt and optional reference images and audio. Inputs longer than 15 seconds are trimmed to the first 15 seconds. Output aspect ratio is explicitly selected from the input display dimensions automatically.
Wan 3.0 · video extend
Wan 3.0 Video Extend continues a video from its final frame and appends a newly generated 2-30 second segment. Inputs longer than 120 seconds retain their last 120 seconds. Supports 480p, 720p, and 1080p output.
flux 3 · image to video draft
Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .
flux 3 · image to video
Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .
flux 3 · start end to video draft
Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .
flux 3 · start end to video
Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .
flux 3 · text to video draft
Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .
flux 3 · text to video
Text-to-image generation with FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene. .
Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.