THE MODEL CATALOG

Chat model APIs & token pricing

Compare text chat models, input and output token prices, context windows and tool support. Use one API key with the OpenAI-compatible chat endpoint.

435 models · Page 1 of 10

google

Google: Gemini 2.5 Flash Lite

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

Chat1,048.576k contextTools
$0.1 input / 1M$0.4 output / 1M
openai

OpenAI: GPT-4.1 Mini

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

Chat1,047.576k contextTools
$0.4 input / 1M$1.6 output / 1M
prism-ml

PrismML: Ternary Bonsai 2 27B

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks...

Chat262.144k contextTools
$0.075 input / 1M$0.5 output / 1M
z-ai

Z.ai: GLM 5.3 FlashX

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Chat1,048.576k contextTools
$0.37 input / 1M$1.25 output / 1M
unbiased

Pareto

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.

Chat262.144k contextTools
$2.5 input / 1M$7.5 output / 1M
~deepseek

DeepSeek: DeepSeek Pro Latest

This model always redirects to the latest model in the DeepSeek Pro family.

Chat1,048.576k contextTools
$0.624 input / 1M$2.88 output / 1M
~deepseek

DeepSeek: DeepSeek Flash Latest

This model always redirects to the latest model in the DeepSeek Flash family.

Chat1,048.576k contextTools
$0.12 input / 1M$0.48 output / 1M
inference-net

Inference.net: Schematron V2 Turbo

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

Chat128k context
$0.03 input / 1M$0.15 output / 1M
inference-net

Inference.net: Schematron V2 Small

Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema...

Chat128k context
$0.05 input / 1M$0.23 output / 1M
~openai

OpenAI: GPT Astra Latest

This model always redirects to the latest model in the GPT Astra family.

Chat1,050k contextTools
$10 input / 1M$50 output / 1M
~openai

OpenAI: GPT Sol Latest

This model always redirects to the latest model in the GPT Sol family.

Chat1,050k contextTools
$2 input / 1M$10 output / 1M
~openai

OpenAI: GPT Terra Latest

This model always redirects to the latest model in the GPT Terra family.

Chat1,050k contextTools
$2 input / 1M$12 output / 1M
~openai

OpenAI: GPT Luna Latest

This model always redirects to the latest model in the GPT Luna family.

Chat1,050k contextTools
$0.2 input / 1M$1.2 output / 1M
sakana

Sakana: Fugu Ultra v2

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...

Chat1,000k contextTools
$5 input / 1M$30 output / 1M
sakana

Sakana: Fugu Max

Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...

Chat1,000k contextTools
$2 input / 1M$6 output / 1M
inclusionai

inclusionAI: Ling 3.0 Flash VL

Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...

Chat131.072k contextTools
$0.06 input / 1M$0.18 output / 1M
inclusionai

inclusionAI: Ling 3.0 Flash VL (free)

Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
deepseek

DeepSeek: DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

Chat1,048.576k contextTools
$0.3 input / 1M$1.2 output / 1M
inception

Inception: Mercury 2.5

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

Chat260k contextTools
$0.04 input / 1M$0.15 output / 1M
nex-agi

Nex AGI: Nex-N2.5-Mini (free)

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
nex-agi

Nex AGI: Nex-N2.5-Pro (free)

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
openai

OpenAI: GPT-6 Astra

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

Chat1,050k contextTools
$10 input / 1M$50 output / 1M
openai

OpenAI: GPT-6 Astra (batch)

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

Chat1,050k contextTools
$5 input / 1M$25 output / 1M
openai

OpenAI: GPT-6 Astra Pro

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Chat1,050k contextTools
$10 input / 1M$50 output / 1M
openai

OpenAI: GPT-6 Astra Pro (batch)

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Chat1,050k contextTools
$5 input / 1M$25 output / 1M
inclusionai

inclusionAI: Ling 3.0 Flash Sante (free)

Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
qwen

Qwen: Qwen3.8 Max (0902)

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

Chat1,000k contextTools
$2 input / 1M$6 output / 1M
meta

Meta: Muse Spark 1.3 Contributor

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

Chat1,048.576k contextTools
$0.1 input / 1M$0.2 output / 1M
meta

Meta: Muse Spark 1.3

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...

Chat1,048.576k contextTools
$1.25 input / 1M$4.25 output / 1M
google

Google: Gemini 3.8 Flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Chat1,048.576k contextTools
$0.75 input / 1M$3.75 output / 1M
google

Google: Gemini 3.8 Flash (batch)

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Chat1,048.576k contextTools
$0.375 input / 1M$1.875 output / 1M
anthropic

Anthropic: Claude Fable 5.1

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Chat1,000k contextTools
$10 input / 1M$50 output / 1M
anthropic

Anthropic: Claude Fable 5.1 (batch)

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Chat1,000k contextTools
$5 input / 1M$25 output / 1M
ibm-granite

IBM: Granite 4.2 8B

Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...

Chat131.072k contextTools
$0.06 input / 1M$0.25 output / 1M
tencent

Tencent: Hy4 preview

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...

Chat1,048.576k contextTools
$0.834 input / 1M$2.501 output / 1M
inclusionai

inclusionAI: Ling 3.0 Flash Fin

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

Chat262.144k contextTools
$0.06 input / 1M$0.18 output / 1M
inclusionai

inclusionAI: Ling 3.0 Flash Fin (free)

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
~z-ai

Z.ai: GLM Flash Latest

This model always redirects to the latest model in the GLM Flash family.

Chat1,310.72k contextTools
$0.075 input / 1M$0.25 output / 1M
qwen

Qwen: Qwen3.8 Flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Chat1,000k contextTools
$0.15 input / 1M$0.47 output / 1M
z-ai

Z.ai: GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Chat1,310.72k contextTools
$0.15 input / 1M$0.5 output / 1M
z-ai

Z.ai: GLM 5.3 Flash (batch)

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Chat1,048.576k contextTools
$0.075 input / 1M$0.25 output / 1M
meta

Meta: Muse Spark 1.2 Contributor

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...

Chat1,048.576k contextTools
$0.1 input / 1M$0.2 output / 1M
deepseek

DeepSeek: DeepSeek V4 Flash Vision Exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

Chat1,048.576k contextTools
$0.22 input / 1M$0.66 output / 1M
deepseek

DeepSeek: DeepSeek V4 Flash Vision Exp (batch)

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

Chat1,048.576k contextTools
$0.11 input / 1M$0.33 output / 1M
tencent

Tencent: Hy-MT2-1.8B

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided...

Chat8.192k context
$0.044 input / 1M$0.177 output / 1M
tencent

Tencent: Hy-MT2-30B-A3B

Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and...

Chat8.192k context
$0.074 input / 1M$0.295 output / 1M
~z-ai

Z.ai: GLM Latest

This model always redirects to the latest GLM model from Z.ai.

Chat1,310.72k contextTools
$0.6545 input / 1M$2.057 output / 1M
tencent

Tencent: Hy-MT2-7B

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.

Chat8.192k context
$0.074 input / 1M$0.295 output / 1M

Model integration notes

Practical prompts, parameter checks and related models. These are starting points for your evaluation, not a performance ranking.

Use one model or an evaluated route.

Automatic mode chooses from your allowed, evaluated models. Catalog inclusion alone does not add a model to that routing pool.

Learn about routing →

Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.