All video models

Jevrouter Video API · video-to-video

ai video editor · video captioner

jev/ai-video-editor/video-captioner:video

Animated word-by-word video captions in about 100 languages: 12 caption styles, AI keyword highlights, emoji, silence and filler removal, translation and SRT subtitle files. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Create an API key

Options and pricing

Task: video-to-video. Catalog base price: $0.08. The exact quote depends on your inputs; the catalog price is not a fixed per-task charge.

We reserve the quoted maximum before submission, settle the actual generation cost and release the remainder. Set jev.max_cost_usd to your maximum reservation. A quote above this limit is rejected before generation.

Input fieldTypeRequiredOptions / default
fontstringNo["default","Montserrat-Black","Montserrat-ExtraBold","Poppins-Black","Poppins-Bold","Inter-ExtraBold","Inter-SemiBold","Anton-Regular","BebasNeue-Regular","Bangers-Regular","LuckiestGuy-Regular","ArchivoBlack-Regular","NotoSans-Black"]
emojibooleanNotrue
videostringYesSee schema
shadowstringNo["default","none","soft","hard","glow"]
strokestringNo["default","none","thin","medium","thick"]
languagestringNo["auto","en","zh-CN","zh-TW","yue","es","fr","de","it","pt","ru","ja","ko","ar","hi","tr","vi","th","id","ms","nl","pl","uk","sv","fi","da","no","nn","cs","sk","ro","hu","el","bg","hr","sr","sl","bs","mk","sq","lt","lv","et","he","fa","ur","ps","sd","bn","as","pa","gu","mr","ne","sa","ta","te","kn","ml","si","my","km","lo","bo","ka","hy","az","kk","uz","tg","tk","mn","ba","tt","be","is","fo","cy","br","eu","gl","ca","oc","lb","la","mt","af","sw","so","am","ha","yo","ln","sn","mg","tl","jw","su","haw","mi","ht","yi"]
positionintegerNoSee schema
templatestringNo["bold-pop","box","karaoke","one-word","word-by-word","minimal","focus","neon","comic","headline","bar","playful"]
font_sizenumberNo1
text_casestringNo["default","upper","original"]
dictionaryarrayNoSee schema
text_colorstringNoSee schema
transcriptstringNoSee schema
aspect_ratiostringNo["original","9:16","16:9","1:1","4:5","5:4","4:3","3:4","3:2","2:3","21:9"]
stroke_colorstringNoSee schema
translate_tostringNo["none","en","zh-CN","zh-TW","yue","es","fr","de","it","pt","ru","ja","ko","ar","hi","tr","vi","th","id","ms","nl","pl","uk","sv","fi","da","no","nn","cs","sk","ro","hu","el","bg","hr","sr","sl","bs","mk","sq","lt","lv","et","he","fa","ur","ps","sd","bn","as","pa","gu","mr","ne","sa","ta","te","kn","ml","si","my","km","lo","bo","ka","hy","az","kk","uz","tg","tk","mn","ba","tt","be","is","fo","cy","br","eu","gl","ca","oc","lb","la","mt","af","sw","so","am","ha","yo","ln","sn","mg","tl","jw","su","haw","mi","ht","yi"]
keyword_colorstringNoSee schema
remove_silencebooleanNofalse
highlight_colorstringNoSee schema
words_per_screenintegerNoSee schema
highlight_keywordsbooleanNotrue
remove_filler_wordsbooleanNofalse
Complete input schema
{
  "type": "object",
  "required": [
    "video"
  ],
  "properties": {
    "font": {
      "enum": [
        "default",
        "Montserrat-Black",
        "Montserrat-ExtraBold",
        "Poppins-Black",
        "Poppins-Bold",
        "Inter-ExtraBold",
        "Inter-SemiBold",
        "Anton-Regular",
        "BebasNeue-Regular",
        "Bangers-Regular",
        "LuckiestGuy-Regular",
        "ArchivoBlack-Regular",
        "NotoSans-Black"
      ],
      "type": "string",
      "default": "default",
      "x-hidden": true,
      "description": "Font family. default uses the template's font.",
      "x-ui-component": "select"
    },
    "emoji": {
      "type": "boolean",
      "default": true,
      "description": "Add animated emoji that match what is said."
    },
    "video": {
      "type": "string",
      "description": "The video to caption. Up to 2 hours is processed."
    },
    "shadow": {
      "enum": [
        "default",
        "none",
        "soft",
        "hard",
        "glow"
      ],
      "type": "string",
      "default": "default",
      "x-hidden": true,
      "description": "Text shadow.",
      "x-ui-component": "select"
    },
    "stroke": {
      "enum": [
        "default",
        "none",
        "thin",
        "medium",
        "thick"
      ],
      "type": "string",
      "default": "default",
      "x-hidden": true,
      "description": "Outline around the letters.",
      "x-ui-component": "select"
    },
    "language": {
      "enum": [
        "auto",
        "en",
        "zh-CN",
        "zh-TW",
        "yue",
        "es",
        "fr",
        "de",
        "it",
        "pt",
        "ru",
        "ja",
        "ko",
        "ar",
        "hi",
        "tr",
        "vi",
        "th",
        "id",
        "ms",
        "nl",
        "pl",
        "uk",
        "sv",
        "fi",
        "da",
        "no",
        "nn",
        "cs",
        "sk",
        "ro",
        "hu",
        "el",
        "bg",
        "hr",
        "sr",
        "sl",
        "bs",
        "mk",
        "sq",
        "lt",
        "lv",
        "et",
        "he",
        "fa",
        "ur",
        "ps",
        "sd",
        "bn",
        "as",
        "pa",
        "gu",
        "mr",
        "ne",
        "sa",
        "ta",
        "te",
        "kn",
        "ml",
        "si",
        "my",
        "km",
        "lo",
        "bo",
        "ka",
        "hy",
        "az",
        "kk",
        "uz",
        "tg",
        "tk",
        "mn",
        "ba",
        "tt",
        "be",
        "is",
        "fo",
        "cy",
        "br",
        "eu",
        "gl",
        "ca",
        "oc",
        "lb",
        "la",
        "mt",
        "af",
        "sw",
        "so",
        "am",
        "ha",
        "yo",
        "ln",
        "sn",
        "mg",
        "tl",
        "jw",
        "su",
        "haw",
        "mi",
        "ht",
        "yi"
      ],
      "type": "string",
      "default": "auto",
      "description": "Language spoken in the video. \"auto\" detects it; set it when detection is unreliable, for example with heavy accents or background music.",
      "x-ui-component": "select"
    },
    "position": {
      "type": "integer",
      "maximum": 100,
      "minimum": 0,
      "x-hidden": true,
      "description": "Vertical centre of the captions as a percentage of the frame height from the top. Leave it empty to use the template's position."
    },
    "template": {
      "enum": [
        "bold-pop",
        "box",
        "karaoke",
        "one-word",
        "word-by-word",
        "minimal",
        "focus",
        "neon",
        "comic",
        "headline",
        "bar",
        "playful"
      ],
      "type": "string",
      "default": "bold-pop",
      "description": "Caption look and animation: bold-pop (bold outline, spoken word pops), box (spoken word on a colour block), karaoke (words fill as spoken), one-word (one big word at a time), word-by-word (words appear as spoken), minimal (clean subtitles), focus (unspoken words dimmed), neon (glow), comic, headline (condensed type), bar (text on a translucent bar), playful.",
      "x-ui-component": "select"
    },
    "font_size": {
      "step": 0.1,
      "type": "number",
      "default": 1,
      "maximum": 2,
      "minimum": 0.5,
      "x-hidden": true,
      "description": "Text size relative to the template, from 0.5 to 2.0.",
      "x-ui-component": "slider"
    },
    "text_case": {
      "enum": [
        "default",
        "upper",
        "original"
      ],
      "type": "string",
      "default": "default",
      "x-hidden": true,
      "description": "upper turns captions into capitals; original keeps the spoken casing.",
      "x-ui-component": "select"
    },
    "dictionary": {
      "type": "array",
      "items": {
        "type": "string"
      },
      "maxItems": 100,
      "x-hidden": true,
      "description": "Names and terms to spell exactly, such as brands or people.",
      "x-placeholder": "e.g. WaveSpeedAI",
      "x-ui-component": "array"
    },
    "text_color": {
      "type": "string",
      "x-hidden": true,
      "description": "Text colour as #RRGGBB or #RRGGBBAA. Empty uses the template."
    },
    "transcript": {
      "type": "string",
      "x-hidden": true,
      "description": "The script of what is said, as plain text in the spoken language (no timestamps or speaker labels). Captions use its wording; timing comes from the audio. Ignored if most of it does not match the speech.",
      "x-ui-component": "textarea"
    },
    "aspect_ratio": {
      "enum": [
        "original",
        "9:16",
        "16:9",
        "1:1",
        "4:5",
        "5:4",
        "4:3",
        "3:4",
        "3:2",
        "2:3",
        "21:9"
      ],
      "type": "string",
      "default": "original",
      "description": "Output width:height. original keeps the source; other ratios keep the whole frame on a blurred fill.",
      "x-ui-component": "select"
    },
    "stroke_color": {
      "type": "string",
      "x-hidden": true,
      "description": "Outline colour. Empty uses the template."
    },
    "translate_to": {
      "enum": [
        "none",
        "en",
        "zh-CN",
        "zh-TW",
        "yue",
        "es",
        "fr",
        "de",
        "it",
        "pt",
        "ru",
        "ja",
        "ko",
        "ar",
        "hi",
        "tr",
        "vi",
        "th",
        "id",
        "ms",
        "nl",
        "pl",
        "uk",
        "sv",
        "fi",
        "da",
        "no",
        "nn",
        "cs",
        "sk",
        "ro",
        "hu",
        "el",
        "bg",
        "hr",
        "sr",
        "sl",
        "bs",
        "mk",
        "sq",
        "lt",
        "lv",
        "et",
        "he",
        "fa",
        "ur",
        "ps",
        "sd",
        "bn",
        "as",
        "pa",
        "gu",
        "mr",
        "ne",
        "sa",
        "ta",
        "te",
        "kn",
        "ml",
        "si",
        "my",
        "km",
        "lo",
        "bo",
        "ka",
        "hy",
        "az",
        "kk",
        "uz",
        "tg",
        "tk",
        "mn",
        "ba",
        "tt",
        "be",
        "is",
        "fo",
        "cy",
        "br",
        "eu",
        "gl",
        "ca",
        "oc",
        "lb",
        "la",
        "mt",
        "af",
        "sw",
        "so",
        "am",
        "ha",
        "yo",
        "ln",
        "sn",
        "mg",
        "tl",
        "jw",
        "su",
        "haw",
        "mi",
        "ht",
        "yi"
      ],
      "type": "string",
      "default": "none",
      "description": "Translate the captions into this language. none keeps the spoken language.",
      "x-ui-component": "select"
    },
    "keyword_color": {
      "type": "string",
      "x-hidden": true,
      "description": "Colour of highlighted keywords. Empty uses the template."
    },
    "remove_silence": {
      "type": "boolean",
      "default": false,
      "description": "Cut pauses between sentences."
    },
    "highlight_color": {
      "type": "string",
      "x-hidden": true,
      "description": "Colour of the word being spoken (or of its box). Empty uses the template."
    },
    "words_per_screen": {
      "type": "integer",
      "maximum": 12,
      "minimum": 1,
      "x-hidden": true,
      "description": "Most words shown at once, 1 to 12. Leave it empty to use the template."
    },
    "highlight_keywords": {
      "type": "boolean",
      "default": true,
      "description": "Colour the most important words."
    },
    "remove_filler_words": {
      "type": "boolean",
      "default": false,
      "description": "Cut hesitations such as um and uh."
    }
  },
  "x-order-properties": [
    "video",
    "template",
    "language",
    "translate_to",
    "highlight_keywords",
    "emoji",
    "remove_silence",
    "remove_filler_words",
    "aspect_ratio",
    "font",
    "font_size",
    "text_color",
    "highlight_color",
    "keyword_color",
    "stroke",
    "stroke_color",
    "shadow",
    "position",
    "text_case",
    "words_per_screen",
    "dictionary",
    "transcript"
  ],
  "additionalProperties": false
}

Required fields and conditional rules come from the published schema. Catalog support does not mean every option has been independently tested by Jevrouter. Synchronous generation, base64 output and customer callbacks are unavailable.

1. Submit a video task

Allow this model on your Jevrouter key. Replace every placeholder below with a value of the correct type, set a USD limit, and supply any conditional fields required by the schema. Media inputs must use your own accessible URLs; this endpoint does not upload files.

curl https://api.jevrouter.io/v1/videos \
  -H "Authorization: Bearer $JEVROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-demo-001" \
  -d @- <<'JSON'
{
  "model": "jev/ai-video-editor/video-captioner:video",
  "input": {
    "font": "default",
    "emoji": true,
    "video": "<video: provide string>",
    "shadow": "default",
    "stroke": "default",
    "language": "auto",
    "template": "bold-pop",
    "font_size": 1,
    "text_case": "default",
    "aspect_ratio": "original",
    "translate_to": "none",
    "remove_silence": false,
    "highlight_keywords": true,
    "remove_filler_words": false
  },
  "jev": {
    "max_cost_usd": "YOUR_USD_LIMIT"
  }
}
JSON

Save data.id and data.urls.get. An accepted task continues even if you disconnect.

Let Jevrouter choose

Use jev/auto as the model and set task to video-to-video. Your project must enable Auto, and both the project and key must allow Auto and the candidate models.

Auto checks the task type and all supplied input fields against each candidate. It compares live quotes for up to eight compatible candidates, shortlisted by catalog base price. It selects the lowest quote within that shortlist and your budget; this is not a quality ranking or a guarantee of the cheapest model in the entire catalog.

Use common options such as duration, resolution and aspect_ratio at the top level where available. Model-specific options belong in input and may narrow the pool. Unsupported options are rejected, never silently dropped. Without task, Auto defaults to text-to-video.

2. Check the result

curl https://api.jevrouter.io/v1/videos/REQUEST_ID \
  -H "Authorization: Bearer $JEVROUTER_API_KEY"

Replace REQUEST_ID with your task ID. Wait at least three seconds between checks. Status progresses from queued to processing, then completed or failed. Results appear in data.outputs as media URLs or structured output objects.

If the submission response is lost, repeat the original request with the same idempotency key. A new key may create another billed task.

Results and billing

View the request in your dashboard or query /v1/requests/REQUEST_ID for reservation and settlement details. Results may be ready before the bill arrives.

Result links are retained for up to seven days and may expire earlier. Save your video while it is available. Anyone with a media URL may be able to access it. Cancellation and customer webhooks are not available.

Build the complete flow