THE MODEL CATALOG
Video model APIs & pricing
Find text-to-video, image-to-video and video processing APIs by task. Compare Wan versions, inspect each model’s parameters and submit asynchronous jobs with one Jevrouter key.
532 models · Page 4 of 12
veo3 · image to video
Google Veo 3 is Google's flagship image-to-video model that creates audio-enabled videos from images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
veo3
Google Veo3 is Google's flagship text-to-video model with built-in audio, producing synchronized video and sound from text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
avatar v · digital twin
HeyGen Avatar V Digital Twin generates natural avatar videos from text or audio with lip-sync, optional captions, background removal, and MP4/WebM output.
video translate
HeyGen Video Translate: translate videos into 70+ languages and 175+ dialects with voice preservation and lip sync. Choose Speed at $0.04/sec or Precision at $0.08/sec. Supports videos up to 120 seconds.
ai dog selfie video
wavespeed-ai/ai-dog-selfie-video
ai kissing
wavespeed-ai/ai-kissing
ai parkour video
wavespeed-ai/ai-parkour-video
ai sketch to video
wavespeed-ai/ai-sketch-to-video
ai talking photos
wavespeed-ai/ai-talking-photos
ai twerk
wavespeed-ai/ai-twerk
ai video ads
wavespeed-ai/ai-video-ads
ai video editor · auto clip
AI Video Editor Auto Clip cuts long or raw footage down to a short highlight clip: it reviews one or several videos end to end, picks the moments that matter, follows your prompt for topic, style and length, and delivers a captioned clip of roughly 15 to 60 seconds. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ai video editor · talking head
AI Video Editor Talking Head cleans up a talking-to-camera video: it removes filler words, false starts, repeats and dead air, keeps the speaker framed, and burns in accurate word-timed captions. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ai video editor · video captioner
Animated word-by-word video captions in about 100 languages: 12 caption styles, AI keyword highlights, emoji, silence and filler removal, translation and SRT subtitle files. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ai virtual outfit tryon
wavespeed-ai/ai-virtual-outfit-tryon
cinematic video generator
WaveSpeed Cinematic Video Generator creates Hollywood-grade videos from text prompts and optional reference images with native audio, director-level camera control, and real-world physics. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
cosmos predict 2.5 · image to video
Cosmos Predict 2.5 Image-to-Video generates video from an image and text prompt using NVIDIA's 2B Cosmos Post-Trained Model. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
cosmos predict 2.5 · text to video
Cosmos Predict 2.5 Text-to-Video generates video from text prompts using NVIDIA's 2B Cosmos Post-Trained Model. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
davinci magihuman · image to video
daVinci MagiHuman is a 15B parameter omni video generation model — the new open-source king on par with WAN 2.5. Generates high-quality AI videos from reference images with optional audio input. Supports digital humans, talking heads, and general video generation.
davinci magihuman · text to video
daVinci MagiHuman is a 15B parameter omni video generation model — the new open-source king on par with WAN 2.5. Generates high-quality AI videos from text prompts with optional audio input. Supports digital humans, talking heads, and general video generation.
depth anything v3 · video
Depth Anything V3 Video turns any video into a temporally consistent depth map video. Depth stays stable across the whole clip with no flicker or brightness pumping, ideal for replicating camera moves and motion with depth-controlled video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
depth anything · video
Depth Anything Video estimates temporally consistent depth maps from video input, with stable depth across the whole clip and no flicker. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
face enhancer · video
Video Face Enhancer restores and sharpens up to three faces in every frame of a video, with stable, flicker-free results and the original audio kept. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
flashvsr
FlashVSR is a fast, high-quality video upscaler that boosts resolution and restores clarity for low-resolution or blurry footage. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
framepack
Framepack is an efficient autoregressive Image-to-Video model that generates smooth, temporally consistent videos from a single image. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ghibli filter · video
wavespeed-ai/ai-ghibli-filter-video
hunyuan avatar
Hunyuan Avatar creates audio-driven talking or singing videos from one image + audio, in 480p/720p up to 120s (starts at $0.15/5s). Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
hunyuan video 1.5 · image to video
HunyuanVideo-1.5 (i2v) is a lightweight 8.3B parameter image-to-video model that generates high-quality videos from images with top-tier visual quality and motion coherence. Optimized for fast inference on consumer-grade GPUs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
hunyuan video 1.5 · text to video
HunyuanVideo-1.5 (t2v) is a lightweight 8.3B parameter text-to-video model that generates high-quality videos with top-tier visual quality and motion coherence. Optimized for fast inference on consumer-grade GPUs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
hunyuan video · Image to Video
Hunyuan i2v turns images and text prompts into high-quality videos, generating coherent short clips from descriptive inputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
hunyuan video · Text to Video
Hunyuan Video (t2v) is an advanced text-to-video model that generates high-quality videos from text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
infinitetalk fast · multi
InfiniteTalk fast multi converts a single image and two audio inputs into multi-character talking or singing videos. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
infinitetalk fast
InfiniteTalk fast converts one photo + audio into audio-driven talking or singing avatar videos (Image-to-Video), up to 10 minutes. Ready-to-use REST API, no coldstarts, affordable pricing.
infinitetalk fast · video to video multi
InfiniteTalk fast video-to-video multi converts a video and two audio inputs into multi-character talking or singing videos. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
infinitetalk fast · video to video
Audio-driven infinitetalk-fast turns one video plus audio into realistic talking or singing videos with lip-sync. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
infinitetalk · multi
InfiniteTalk Multi converts a single image and two audio inputs into multi-character talking or singing videos at up to 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
infinitetalk
InfiniteTalk converts one photo + audio into audio-driven talking or singing avatar videos (Image-to-Video), up to 10 minutes, 720p tier $0.30/5s. Ready-to-use REST API, no coldstarts, affordable pricing.
infinitetalk · video to video multi
InfiniteTalk Video-to-Video Multi converts a video and two audio inputs into multi-character talking or singing videos at up to 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
infinitetalk · video to video
Audio-driven InfiniteTalk turns one video plus audio into realistic talking or singing videos with lip-sync in 480p or 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
kandinsky5 Pro · image to video
Kandinsky 5 Pro Image-to-Video generation model. Transform static images into dynamic 5-second videos with text prompts. Supports 512P and 1024P resolutions. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
kandinsky5 Pro · text to video
Kandinsky 5 Pro Text-to-Video generation model. Create dynamic 5-second videos from text prompts. Supports 512P and 1024P resolutions with multiple aspect ratios. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
latentsync
LatentSync synchronizes video and audio inputs to generate seamless synchronized content. Perfect for lip-syncing, audio dubbing, and video-audio alignment tasks.
longcat avatar 1.5 · multi
LongCat Avatar 1.5 Multi converts a single image and two audio inputs into multi-character talking or singing videos at up to 720p, capped at 64 seconds per clip. Ready-to-use REST API, no coldstarts, affordable pricing.
longcat avatar 1.5
LongCat Avatar 1.5 is the upgraded version of LongCat Avatar with sharper lip sync and faster generation. Converts one photo + audio into audio-driven talking or singing avatar videos (Image-to-Video), capped at 64 seconds per clip. Ready-to-use REST API, no coldstarts, affordable pricing.
longcat avatar
LongCat Avatar produces super-realistic, lip-synchronized long video generation with natural dynamics and consistent identity. Converts one photo + audio into audio-driven talking or singing avatar videos (Image-to-Video), up to 2 minutes, 720p tier $0.30/5s. Ready-to-use REST API, no coldstarts, affordable pricing.
ltx 2 19b · control
LTX-2 19B ControlNet generates synchronized audio-video (up to 20s) from video input with pose, depth, or canny edge guidance. Supports audio preservation, generation, or removal for flexible video transformation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ltx 2 19b · image to video LoRA
LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access. This version supports custom LoRAs for style personalization.
ltx 2 19b · image to video
LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.
Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.