THE MODEL CATALOG

Video model APIs & pricing

Find text-to-video, image-to-video and video processing APIs by task. Compare Wan versions, inspect each model’s parameters and submit asynchronous jobs with one Jevrouter key.

532 models · Page 4 of 12

google

veo3 · image to video

Google Veo 3 is Google's flagship image-to-video model that creates audio-enabled videos from images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
google

veo3

Google Veo3 is Google's flagship text-to-video model with built-in audio, producing synchronized video and sound from text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
heygen

avatar v · digital twin

HeyGen Avatar V Digital Twin generates natural avatar videos from text or audio with lip-sync, optional captions, background removal, and MP4/WebM output.

Video
Per video · quoted before submission
heygen

video translate

HeyGen Video Translate: translate videos into 70+ languages and 175+ dialects with voice preservation and lip sync. Choose Speed at $0.04/sec or Precision at $0.08/sec. Supports videos up to 120 seconds.

Video
Per video · quoted before submission
jev

ai dog selfie video

wavespeed-ai/ai-dog-selfie-video

Video
Per video · quoted before submission
jev

ai kissing

wavespeed-ai/ai-kissing

Video
Per video · quoted before submission
jev

ai parkour video

wavespeed-ai/ai-parkour-video

Video
Per video · quoted before submission
jev

ai sketch to video

wavespeed-ai/ai-sketch-to-video

Video
Per video · quoted before submission
jev

ai talking photos

wavespeed-ai/ai-talking-photos

Video
Per video · quoted before submission
jev

ai twerk

wavespeed-ai/ai-twerk

Video
Per video · quoted before submission
jev

ai video ads

wavespeed-ai/ai-video-ads

Video
Per video · quoted before submission
jev

ai video editor · auto clip

AI Video Editor Auto Clip cuts long or raw footage down to a short highlight clip: it reviews one or several videos end to end, picks the moments that matter, follows your prompt for topic, style and length, and delivers a captioned clip of roughly 15 to 60 seconds. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ai video editor · talking head

AI Video Editor Talking Head cleans up a talking-to-camera video: it removes filler words, false starts, repeats and dead air, keeps the speaker framed, and burns in accurate word-timed captions. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ai video editor · video captioner

Animated word-by-word video captions in about 100 languages: 12 caption styles, AI keyword highlights, emoji, silence and filler removal, translation and SRT subtitle files. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ai virtual outfit tryon

wavespeed-ai/ai-virtual-outfit-tryon

Video
Per video · quoted before submission
jev

cinematic video generator

WaveSpeed Cinematic Video Generator creates Hollywood-grade videos from text prompts and optional reference images with native audio, director-level camera control, and real-world physics. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
jev

cosmos predict 2.5 · image to video

Cosmos Predict 2.5 Image-to-Video generates video from an image and text prompt using NVIDIA's 2B Cosmos Post-Trained Model. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
jev

cosmos predict 2.5 · text to video

Cosmos Predict 2.5 Text-to-Video generates video from text prompts using NVIDIA's 2B Cosmos Post-Trained Model. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
jev

davinci magihuman · image to video

daVinci MagiHuman is a 15B parameter omni video generation model — the new open-source king on par with WAN 2.5. Generates high-quality AI videos from reference images with optional audio input. Supports digital humans, talking heads, and general video generation.

Video
Per video · quoted before submission
jev

davinci magihuman · text to video

daVinci MagiHuman is a 15B parameter omni video generation model — the new open-source king on par with WAN 2.5. Generates high-quality AI videos from text prompts with optional audio input. Supports digital humans, talking heads, and general video generation.

Video
Per video · quoted before submission
jev

depth anything v3 · video

Depth Anything V3 Video turns any video into a temporally consistent depth map video. Depth stays stable across the whole clip with no flicker or brightness pumping, ideal for replicating camera moves and motion with depth-controlled video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

depth anything · video

Depth Anything Video estimates temporally consistent depth maps from video input, with stable depth across the whole clip and no flicker. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

face enhancer · video

Video Face Enhancer restores and sharpens up to three faces in every frame of a video, with stable, flicker-free results and the original audio kept. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

flashvsr

FlashVSR is a fast, high-quality video upscaler that boosts resolution and restores clarity for low-resolution or blurry footage. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

framepack

Framepack is an efficient autoregressive Image-to-Video model that generates smooth, temporally consistent videos from a single image. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ghibli filter · video

wavespeed-ai/ai-ghibli-filter-video

Video
Per video · quoted before submission
jev

hunyuan avatar

Hunyuan Avatar creates audio-driven talking or singing videos from one image + audio, in 480p/720p up to 120s (starts at $0.15/5s). Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

hunyuan video 1.5 · image to video

HunyuanVideo-1.5 (i2v) is a lightweight 8.3B parameter image-to-video model that generates high-quality videos from images with top-tier visual quality and motion coherence. Optimized for fast inference on consumer-grade GPUs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

hunyuan video 1.5 · text to video

HunyuanVideo-1.5 (t2v) is a lightweight 8.3B parameter text-to-video model that generates high-quality videos with top-tier visual quality and motion coherence. Optimized for fast inference on consumer-grade GPUs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

hunyuan video · Image to Video

Hunyuan i2v turns images and text prompts into high-quality videos, generating coherent short clips from descriptive inputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

hunyuan video · Text to Video

Hunyuan Video (t2v) is an advanced text-to-video model that generates high-quality videos from text prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

infinitetalk fast · multi

InfiniteTalk fast multi converts a single image and two audio inputs into multi-character talking or singing videos. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

infinitetalk fast

InfiniteTalk fast converts one photo + audio into audio-driven talking or singing avatar videos (Image-to-Video), up to 10 minutes. Ready-to-use REST API, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

infinitetalk fast · video to video multi

InfiniteTalk fast video-to-video multi converts a video and two audio inputs into multi-character talking or singing videos. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

infinitetalk fast · video to video

Audio-driven infinitetalk-fast turns one video plus audio into realistic talking or singing videos with lip-sync. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

infinitetalk · multi

InfiniteTalk Multi converts a single image and two audio inputs into multi-character talking or singing videos at up to 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

infinitetalk

InfiniteTalk converts one photo + audio into audio-driven talking or singing avatar videos (Image-to-Video), up to 10 minutes, 720p tier $0.30/5s. Ready-to-use REST API, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

infinitetalk · video to video multi

InfiniteTalk Video-to-Video Multi converts a video and two audio inputs into multi-character talking or singing videos at up to 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

infinitetalk · video to video

Audio-driven InfiniteTalk turns one video plus audio into realistic talking or singing videos with lip-sync in 480p or 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

kandinsky5 Pro · image to video

Kandinsky 5 Pro Image-to-Video generation model. Transform static images into dynamic 5-second videos with text prompts. Supports 512P and 1024P resolutions. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
jev

kandinsky5 Pro · text to video

Kandinsky 5 Pro Text-to-Video generation model. Create dynamic 5-second videos from text prompts. Supports 512P and 1024P resolutions with multiple aspect ratios. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Video
Per video · quoted before submission
jev

latentsync

LatentSync synchronizes video and audio inputs to generate seamless synchronized content. Perfect for lip-syncing, audio dubbing, and video-audio alignment tasks.

Video
Per video · quoted before submission
jev

longcat avatar 1.5 · multi

LongCat Avatar 1.5 Multi converts a single image and two audio inputs into multi-character talking or singing videos at up to 720p, capped at 64 seconds per clip. Ready-to-use REST API, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

longcat avatar 1.5

LongCat Avatar 1.5 is the upgraded version of LongCat Avatar with sharper lip sync and faster generation. Converts one photo + audio into audio-driven talking or singing avatar videos (Image-to-Video), capped at 64 seconds per clip. Ready-to-use REST API, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

longcat avatar

LongCat Avatar produces super-realistic, lip-synchronized long video generation with natural dynamics and consistent identity. Converts one photo + audio into audio-driven talking or singing avatar videos (Image-to-Video), up to 2 minutes, 720p tier $0.30/5s. Ready-to-use REST API, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ltx 2 19b · control

LTX-2 19B ControlNet generates synchronized audio-video (up to 20s) from video input with pose, depth, or canny edge guidance. Supports audio preservation, generation, or removal for flexible video transformation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video
Per video · quoted before submission
jev

ltx 2 19b · image to video LoRA

LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access. This version supports custom LoRAs for style personalization.

Video
Per video · quoted before submission
jev

ltx 2 19b · image to video

LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: synchronized audio and video, high fidelity, multiple performance modes, production-ready outputs, API access, and open access.

Video
Per video · quoted before submission

Choose a model. Keep one video endpoint.

Pin a variant or let Auto compare compatible tasks within your key’s model pool and reservation limit.

Video API tutorial

Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.