THE MODEL CATALOG
Video model APIs & pricing
Find text-to-video, image-to-video and video processing APIs by task. Compare Wan versions, inspect each model’s parameters and submit asynchronous jobs with one Jevrouter key.
532 models · Page 5 of 12
ltx 2 19b · lipsync
LTX-2 Lipsync is an audio-driven digital human model that generates synchronized talking head videos from a reference image and audio input. It leverages the LTX-2 19B DiT architecture to produce high-fidelity lip-synced videos with natural head movements.
ltx 2 19b · text to video LoRA
LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access. This version supports custom LoRAs for style personalization.
ltx 2 19b · text to video
LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.
ltx 2 19b · video upscaler
LTX-2 19B Video Upscaler converts low-resolution videos into crisp 4K footage with seamless motion dynamics and frame consistency. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ltx 2.3 · image to video LoRA
LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.
ltx 2.3 · image to video
LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.
ltx 2.3 · lipsync
LTX-2.3 Lipsync is an audio-driven digital human model that generates synchronized talking head videos from a reference image and audio input. It leverages the LTX-2.3 19B DiT architecture to produce high-fidelity lip-synced videos with natural head movements.
ltx 2.3 spicy · image to video LoRA
LTX 2.3 Spicy LoRA converts a reference image and prompt into expressive video with selectable LoRA presets, optional LoRA strength overrides, duration, and resolution. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ltx 2.3 spicy · image to video
LTX 2.3 Spicy converts a reference image and prompt into expressive video with selectable style preset, duration, and resolution. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ltx 2.3 · text to video LoRA
LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.
ltx 2.3 · text to video
LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.
ltx 2.3 · video extend
LTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.
ltx 2.5 · image to video
LTX-2.5 animates a first-frame image into synchronized audio-video content, with optional last-frame guidance and resolutions up to 4K.
ltx 2.5 · text to video
LTX-2.5 generates synchronized audio-video content from text prompts with high-fidelity motion, flexible duration, and resolutions up to 4K.
ltx 2 · video extend
LTX Video 2.0 Pro extends existing videos by generating new content at the start or end. Supports prompt-guided extension up to 20 seconds. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
ltx video v097 · Image to Video 480p
Generate 480p videos from text prompts and images with LTX Video 0.9.7 (i2v-480p). Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ltx video v097 · Image to Video 720p
Generate 720p videos from text prompts and images with LTX Video-0.9.7 for consistent, high-quality image-to-video results. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
lynx
Lynx by ByteDance converts images into subject-consistent videos, preserving visual consistency across frames for seamless animations and reliable subject fidelity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
magi 1 24b
Magi-1 is a cinematic video generation model with strong understanding of physical interactions and cinematic prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
minimax h3 · controlnet union
MiniMax H3 Open Weights ControlNet Union generates a new video that follows the motion and composition of a source video. Pose, depth, edges, lines, scribble or grayscale structure is extracted from the source automatically and guides the output, optionally with reference images for the subject or style, with native stereo audio generated in the same pass. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
minimax h3 · image to video LoRA
Open-weights MiniMax H3 image-to-video on WaveSpeed infrastructure: animate a first-frame image (optionally with a last frame) into a coherent 480p/540p/768p/1080p video with native stereo audio, 5-15 second durations, and affordable per-second pricing, with custom LoRA support.
minimax h3 · image to video spicy
MiniMax H3 Open Weights Spicy (Image-to-Video) generates unlimited high-quality cinematic clips from a reference image and prompt, optimized for scalable content generation with smooth, expressive animations, stable aesthetics, and native stereo audio.
minimax h3 · image to video
Open-weights MiniMax H3 image-to-video on WaveSpeed infrastructure: animate a first-frame image (optionally with a last frame) into a coherent 480p/540p/768p/1080p video with native stereo audio, 5-15 second durations, and affordable per-second pricing.
minimax h3 · reference to video LoRA
Open-weights MiniMax H3 reference-to-video on WaveSpeed infrastructure: generate coherent 480p/540p/768p/1080p videos with native stereo audio guided by up to 9 reference images, 3 reference videos, and 3 reference audios, with custom LoRA support.
minimax h3 · reference to video
Open-weights MiniMax H3 reference-to-video on WaveSpeed infrastructure: generate coherent 480p/540p/768p/1080p videos with native stereo audio guided by up to 9 reference images, 3 reference videos, and 3 reference audios.
minimax h3 singularity · image to video LoRA
Open-weights MiniMax H3 Singularity image-to-video on WaveSpeed infrastructure: animate a first-frame image (optionally with a last frame) into a coherent 480p/540p/768p/1080p video with native stereo audio, 5-15 second durations, and affordable per-second pricing, with custom LoRA support.
minimax h3 singularity · image to video
Open-weights MiniMax H3 Singularity image-to-video on WaveSpeed infrastructure: animate a first-frame image (optionally with a last frame) into a coherent 480p/540p/768p/1080p video with native stereo audio, 5-15 second durations, and affordable per-second pricing.
minimax h3 singularity · reference to video LoRA
Open-weights MiniMax H3 Singularity reference-to-video on WaveSpeed infrastructure: generate coherent 480p/540p/768p/1080p videos with native stereo audio guided by up to 9 reference images, 3 reference videos, and 3 reference audios, with custom LoRA support.
minimax h3 singularity · reference to video
Open-weights MiniMax H3 Singularity reference-to-video on WaveSpeed infrastructure: generate coherent 480p/540p/768p/1080p videos with native stereo audio guided by up to 9 reference images, 3 reference videos, and 3 reference audios.
minimax h3 · text to video LoRA
Open-weights MiniMax H3 text-to-video on WaveSpeed infrastructure: coherent 480p/540p/768p/1080p videos with native stereo audio from a single prompt, 5-15 second durations, flexible aspect ratios, and affordable per-second pricing, with custom LoRA support.
minimax h3 · text to video
Open-weights MiniMax H3 text-to-video on WaveSpeed infrastructure: coherent 480p/540p/768p/1080p videos with native stereo audio from a single prompt, 5-15 second durations, flexible aspect ratios, and affordable per-second pricing.
minimax h3 · video edit
MiniMax H3 Open Weights Video-Edit edits an input video from a natural-language prompt. The input video drives subject identity, composition, and motion while the model rewrites lighting, style, weather, environment, or specific elements as instructed, with native stereo audio generated in the same pass.
minimax h3 · video extend
MiniMax H3 Open Weights Video-Extend appends a new cinematic continuation to an existing video. A fresh segment with native stereo audio is generated from the input video's last frame and a natural-language prompt, then concatenated onto the original.
multitalk
MultiTalk converts one image and audio into audio-driven talking/singing videos (Image-to-Video), supporting up to 10 minutes. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
music video generator
AI Music Video Generator transforms audio + a single photo into a full music video with cinematic camera angles, smooth transitions, and perfect lip sync. Up to 10 minutes, 480p or 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
open video · image to video LoRA
OpenVideo image-to-video (LoRA tier) — full creative control over short cinematic clips with a native audio track, plus an optional preset and per-LoRA strength dict to drive style, motion and look-and-feel. Unlimited generation freedom on the same base model as the standard endpoint. Supports 480p / 720p / 1080p output and 3-20 s duration.
open video · image to video
OpenVideo image-to-video gives you full creative control over short cinematic clips with a native audio track. Unlimited generation freedom — describe any scene, style or motion and the model follows the prompt as-given, from a single reference image. Supports 480p / 720p / 1080p output and 3-20 s duration tiers.
rife
RIFE Video Interpolation generates smooth intermediate frames between existing video frames for higher frame rates and smoother motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
sam3 video rle
A unified foundation model for prompt-based segmentation in video. Returns RLE encoded masks.
sam3 video
A unified foundation model for prompt-based segmentation in video
scail 2
SCAIL-2 is an end-to-end character animation model that preserves identity and motion from a single reference image and a driving video. Built on Wan 2.1 14B with multi-identity color-coded SAM3 masks, supporting both Animation mode (animate reference) and Replacement mode (swap subject), at 480p or 720p.
scail
SCAIL enables high-fidelity character animation using reference images. It handles large motion variations, stylized characters, and multi-character interactions without explicit per-frame structural guidance. Ready-to-use REST inference API, no coldstarts, affordable pricing.
seedvr2 · video
SeedVR2 Video Upscaler turns low-resolution videos into high-fidelity 4K with seamless motion and consistent frames. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
short video generator
WaveSpeed Short Video Generator creates professional short-form videos from text prompts and optional reference images with native audio, smooth motion, and versatile aspect ratios. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
skyreels v3 · talking avatar
19B-parameter multimodal talking avatar generation from portrait and audio with precise lip sync up to 20 seconds at 720p
soulx flashhead
Real-time streaming talking head video generation from portrait image and audio with 96 FPS on RTX 4090
steady dancer
SteadyDancer is a 14B-parameter human image animation framework that transforms static images into coherent dance videos. Features first-frame preservation, robust identity consistency, and temporal coherence for realistic motion generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
tiktok video generator
WaveSpeed TikTok Video Generator creates viral-ready videos from text prompts and optional reference images with native audio, dynamic transitions, and scroll-stopping motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.