THE MODEL CATALOG

Video model APIs & pricing

Find text-to-video, image-to-video and video processing APIs by task. Compare Wan versions, inspect each model’s parameters and submit asynchronous jobs with one Jevrouter key.

532 models · Page 5 of 12

jev

ltx 2 19b · lipsync

LTX-2 Lipsync is an audio-driven digital human model that generates synchronized talking head videos from a reference image and audio input. It leverages the LTX-2 19B DiT architecture to produce high-fidelity lip-synced videos with natural head movements.

Video
Per video · quoted before submission
jev

ltx 2 19b · text to video LoRA

LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access. This version supports custom LoRAs for style personalization.

Video
Per video · quoted before submission
jev

ltx 2 19b · text to video

LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.

Video
Per video · quoted before submission
jev

ltx 2 19b · video upscaler

LTX-2 19B Video Upscaler converts low-resolution videos into crisp 4K footage with seamless motion dynamics and frame consistency. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ltx 2.3 · image to video LoRA

LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.

Video
Per video · quoted before submission
jev

ltx 2.3 · image to video

LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.

Video
Per video · quoted before submission
jev

ltx 2.3 · lipsync

LTX-2.3 Lipsync is an audio-driven digital human model that generates synchronized talking head videos from a reference image and audio input. It leverages the LTX-2.3 19B DiT architecture to produce high-fidelity lip-synced videos with natural head movements.

Video
Per video · quoted before submission
jev

ltx 2.3 spicy · image to video LoRA

LTX 2.3 Spicy LoRA converts a reference image and prompt into expressive video with selectable LoRA presets, optional LoRA strength overrides, duration, and resolution. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ltx 2.3 spicy · image to video

LTX 2.3 Spicy converts a reference image and prompt into expressive video with selectable style preset, duration, and resolution. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ltx 2.3 · text to video LoRA

LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.

Video
Per video · quoted before submission
jev

ltx 2.3 · text to video

LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.

Video
Per video · quoted before submission
jev

ltx 2.3 · video extend

LTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.

Video
Per video · quoted before submission
jev

ltx 2.5 · image to video

LTX-2.5 animates a first-frame image into synchronized audio-video content, with optional last-frame guidance and resolutions up to 4K.

Video
Per video · quoted before submission
jev

ltx 2.5 · text to video

LTX-2.5 generates synchronized audio-video content from text prompts with high-fidelity motion, flexible duration, and resolutions up to 4K.

Video
Per video · quoted before submission
jev

ltx 2 · video extend

LTX Video 2.0 Pro extends existing videos by generating new content at the start or end. Supports prompt-guided extension up to 20 seconds. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
jev

ltx video v097 · Image to Video 480p

Generate 480p videos from text prompts and images with LTX Video 0.9.7 (i2v-480p). Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ltx video v097 · Image to Video 720p

Generate 720p videos from text prompts and images with LTX Video-0.9.7 for consistent, high-quality image-to-video results. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

lynx

Lynx by ByteDance converts images into subject-consistent videos, preserving visual consistency across frames for seamless animations and reliable subject fidelity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

magi 1 24b

Magi-1 is a cinematic video generation model with strong understanding of physical interactions and cinematic prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

minimax h3 · controlnet union

MiniMax H3 Open Weights ControlNet Union generates a new video that follows the motion and composition of a source video. Pose, depth, edges, lines, scribble or grayscale structure is extracted from the source automatically and guides the output, optionally with reference images for the subject or style, with native stereo audio generated in the same pass. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

minimax h3 · image to video LoRA

Open-weights MiniMax H3 image-to-video on WaveSpeed infrastructure: animate a first-frame image (optionally with a last frame) into a coherent 480p/540p/768p/1080p video with native stereo audio, 5-15 second durations, and affordable per-second pricing, with custom LoRA support.

Video
Per video · quoted before submission
jev

minimax h3 · image to video spicy

MiniMax H3 Open Weights Spicy (Image-to-Video) generates unlimited high-quality cinematic clips from a reference image and prompt, optimized for scalable content generation with smooth, expressive animations, stable aesthetics, and native stereo audio.

Video
Per video · quoted before submission
jev

minimax h3 · image to video

Open-weights MiniMax H3 image-to-video on WaveSpeed infrastructure: animate a first-frame image (optionally with a last frame) into a coherent 480p/540p/768p/1080p video with native stereo audio, 5-15 second durations, and affordable per-second pricing.

Video
Per video · quoted before submission
jev

minimax h3 · reference to video LoRA

Open-weights MiniMax H3 reference-to-video on WaveSpeed infrastructure: generate coherent 480p/540p/768p/1080p videos with native stereo audio guided by up to 9 reference images, 3 reference videos, and 3 reference audios, with custom LoRA support.

Video
Per video · quoted before submission
jev

minimax h3 · reference to video

Open-weights MiniMax H3 reference-to-video on WaveSpeed infrastructure: generate coherent 480p/540p/768p/1080p videos with native stereo audio guided by up to 9 reference images, 3 reference videos, and 3 reference audios.

Video
Per video · quoted before submission
jev

minimax h3 singularity · image to video LoRA

Open-weights MiniMax H3 Singularity image-to-video on WaveSpeed infrastructure: animate a first-frame image (optionally with a last frame) into a coherent 480p/540p/768p/1080p video with native stereo audio, 5-15 second durations, and affordable per-second pricing, with custom LoRA support.

Video
Per video · quoted before submission
jev

minimax h3 singularity · image to video

Open-weights MiniMax H3 Singularity image-to-video on WaveSpeed infrastructure: animate a first-frame image (optionally with a last frame) into a coherent 480p/540p/768p/1080p video with native stereo audio, 5-15 second durations, and affordable per-second pricing.

Video
Per video · quoted before submission
jev

minimax h3 singularity · reference to video LoRA

Open-weights MiniMax H3 Singularity reference-to-video on WaveSpeed infrastructure: generate coherent 480p/540p/768p/1080p videos with native stereo audio guided by up to 9 reference images, 3 reference videos, and 3 reference audios, with custom LoRA support.

Video
Per video · quoted before submission
jev

minimax h3 singularity · reference to video

Open-weights MiniMax H3 Singularity reference-to-video on WaveSpeed infrastructure: generate coherent 480p/540p/768p/1080p videos with native stereo audio guided by up to 9 reference images, 3 reference videos, and 3 reference audios.

Video
Per video · quoted before submission
jev

minimax h3 · text to video LoRA

Open-weights MiniMax H3 text-to-video on WaveSpeed infrastructure: coherent 480p/540p/768p/1080p videos with native stereo audio from a single prompt, 5-15 second durations, flexible aspect ratios, and affordable per-second pricing, with custom LoRA support.

Video
Per video · quoted before submission
jev

minimax h3 · text to video

Open-weights MiniMax H3 text-to-video on WaveSpeed infrastructure: coherent 480p/540p/768p/1080p videos with native stereo audio from a single prompt, 5-15 second durations, flexible aspect ratios, and affordable per-second pricing.

Video
Per video · quoted before submission
jev

minimax h3 · video edit

MiniMax H3 Open Weights Video-Edit edits an input video from a natural-language prompt. The input video drives subject identity, composition, and motion while the model rewrites lighting, style, weather, environment, or specific elements as instructed, with native stereo audio generated in the same pass.

Video
Per video · quoted before submission
jev

minimax h3 · video extend

MiniMax H3 Open Weights Video-Extend appends a new cinematic continuation to an existing video. A fresh segment with native stereo audio is generated from the input video's last frame and a natural-language prompt, then concatenated onto the original.

Video
Per video · quoted before submission
jev

multitalk

MultiTalk converts one image and audio into audio-driven talking/singing videos (Image-to-Video), supporting up to 10 minutes. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

music video generator

AI Music Video Generator transforms audio + a single photo into a full music video with cinematic camera angles, smooth transitions, and perfect lip sync. Up to 10 minutes, 480p or 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

open video · image to video LoRA

OpenVideo image-to-video (LoRA tier) — full creative control over short cinematic clips with a native audio track, plus an optional preset and per-LoRA strength dict to drive style, motion and look-and-feel. Unlimited generation freedom on the same base model as the standard endpoint. Supports 480p / 720p / 1080p output and 3-20 s duration.

Video
Per video · quoted before submission
jev

open video · image to video

OpenVideo image-to-video gives you full creative control over short cinematic clips with a native audio track. Unlimited generation freedom — describe any scene, style or motion and the model follows the prompt as-given, from a single reference image. Supports 480p / 720p / 1080p output and 3-20 s duration tiers.

Video
Per video · quoted before submission
jev

rife

RIFE Video Interpolation generates smooth intermediate frames between existing video frames for higher frame rates and smoother motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
jev

sam3 video rle

A unified foundation model for prompt-based segmentation in video. Returns RLE encoded masks.

Video
Per video · quoted before submission
jev

sam3 video

A unified foundation model for prompt-based segmentation in video

Video
Per video · quoted before submission
jev

scail 2

SCAIL-2 is an end-to-end character animation model that preserves identity and motion from a single reference image and a driving video. Built on Wan 2.1 14B with multi-identity color-coded SAM3 masks, supporting both Animation mode (animate reference) and Replacement mode (swap subject), at 480p or 720p.

Video
Per video · quoted before submission
jev

scail

SCAIL enables high-fidelity character animation using reference images. It handles large motion variations, stylized characters, and multi-character interactions without explicit per-frame structural guidance. Ready-to-use REST inference API, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

seedvr2 · video

SeedVR2 Video Upscaler turns low-resolution videos into high-fidelity 4K with seamless motion and consistent frames. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

short video generator

WaveSpeed Short Video Generator creates professional short-form videos from text prompts and optional reference images with native audio, smooth motion, and versatile aspect ratios. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
jev

skyreels v3 · talking avatar

19B-parameter multimodal talking avatar generation from portrait and audio with precise lip sync up to 20 seconds at 720p

Video
Per video · quoted before submission
jev

soulx flashhead

Real-time streaming talking head video generation from portrait image and audio with 96 FPS on RTX 4090

Video
Per video · quoted before submission
jev

steady dancer

SteadyDancer is a 14B-parameter human image animation framework that transforms static images into coherent dance videos. Features first-frame preservation, robust identity consistency, and temporal coherence for realistic motion generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

tiktok video generator

WaveSpeed TikTok Video Generator creates viral-ready videos from text prompts and optional reference images with native audio, dynamic transitions, and scroll-stopping motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission

Choose a model. Keep one video endpoint.

Pin a variant or let Auto compare compatible tasks within your key’s model pool and reservation limit.

Video API tutorial

Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.