THE MODEL CATALOG

Chat model APIs & token pricing

Compare text chat models, input and output token prices, context windows and tool support. Use one API key with the OpenAI-compatible chat endpoint.

435 models · Page 3 of 10

openai

OpenAI: GPT-5.6 Terra Pro (batch)

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Chat1,050k contextTools
$1 input / 1M$6 output / 1M
openai

OpenAI: GPT-5.6 Terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Chat1,050k contextTools
$2 input / 1M$12 output / 1M
openai

OpenAI: GPT-5.6 Terra (batch)

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Chat1,050k contextTools
$1 input / 1M$6 output / 1M
openai

OpenAI: GPT-5.6 Sol Pro

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Chat1,050k contextTools
$2 input / 1M$10 output / 1M
openai

OpenAI: GPT-5.6 Sol Pro (batch)

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Chat1,050k contextTools
$1 input / 1M$5 output / 1M
openai

OpenAI: GPT-5.6 Sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Chat1,050k contextTools
$2 input / 1M$10 output / 1M
openai

OpenAI: GPT-5.6 Sol (batch)

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Chat1,050k contextTools
$1 input / 1M$5 output / 1M
x-ai

SpaceXAI: Grok 4.5

Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.

Chat500k contextTools
$2 input / 1M$6 output / 1M
~x-ai

xAI: Grok Latest

This model always redirects to the latest Grok model from xAI.

Chat500k contextTools
$1.6 input / 1M$4.8 output / 1M
aion-labs

AionLabs: Aion-3.0-Mini

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...

Chat131.072k contextTools
$0.7 input / 1M$1.4 output / 1M
aion-labs

AionLabs: Aion-3.0

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...

Chat131.072k contextTools
$3 input / 1M$6 output / 1M
tencent

Tencent: Hy3

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

Chat262.144k contextTools
$0.132 input / 1M$0.528 output / 1M
poolside

Poolside: Laguna XS 2.1

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

Chat262.144k contextTools
$0.06 input / 1M$0.12 output / 1M
poolside

Poolside: Laguna XS 2.1 (free)

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
anthropic

Anthropic: Claude Sonnet 5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Chat1,000k contextTools
$2 input / 1M$10 output / 1M
anthropic

Anthropic: Claude Sonnet 5 (batch)

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Chat1,000k contextTools
$1 input / 1M$5 output / 1M
sakana

Sakana: Fugu Ultra

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...

Chat1,000k contextTools
$5 input / 1M$30 output / 1M
cohere

Cohere: North Mini Code (free)

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...

Chat256k contextTools
$0 input / 1M$0 output / 1M
z-ai

Z.ai: GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Chat1,048.576k contextTools
$0.6496 input / 1M$2.0416 output / 1M
z-ai

Z.ai: GLM 5.2 (batch)

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Chat1,048.576k contextTools
$0.7 input / 1M$2.2 output / 1M
z-ai

Z.ai: GLM 5.2 (free)

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Chat32.768k context
$0 input / 1M$0 output / 1M
openrouter

OpenRouter: Fusion

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...

Chat1,000k context
Price depends on the selected model
moonshotai

MoonshotAI: Kimi K2.7 Code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

Chat262.144k contextTools
$0.7062 input / 1M$3.21 output / 1M
~anthropic

Anthropic: Claude Fable Latest

This model always redirects to the latest model in the Claude Fable family.

Chat1,000k contextTools
$10 input / 1M$50 output / 1M
anthropic

Anthropic: Claude Fable 5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Chat1,000k contextTools
$10 input / 1M$50 output / 1M
anthropic

Anthropic: Claude Fable 5 (batch)

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Chat1,000k contextTools
$5 input / 1M$25 output / 1M
nvidia

NVIDIA: Nemotron 3.5 Content Safety

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

Chat131.072k context
$0.2 input / 1M$0.2 output / 1M
nvidia

NVIDIA: Nemotron 3.5 Content Safety (free)

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

Chat128k context
$0 input / 1M$0 output / 1M
nvidia

NVIDIA: Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Chat262.144k contextTools
$0.6 input / 1M$2.4 output / 1M
nvidia

NVIDIA: Nemotron 3 Ultra (free)

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Chat1,000k contextTools
$0 input / 1M$0 output / 1M
qwen

Qwen: Qwen3.7 Plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

Chat1,000k contextTools
$0.32 input / 1M$1.28 output / 1M
minimax

MiniMax: MiniMax M3

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Chat1,048.576k contextTools
$0.3 input / 1M$1.2 output / 1M
stepfun

StepFun: Step 3.7 Flash

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

Chat262.144k contextTools
$0.2 input / 1M$1.15 output / 1M
anthropic

Anthropic: Claude Opus 4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

Chat1,000k contextTools
$5 input / 1M$25 output / 1M
anthropic

Anthropic: Claude Opus 4.8 (batch)

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

Chat1,000k contextTools
$2.5 input / 1M$12.5 output / 1M
qwen

Qwen: Qwen3.7 Max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

Chat1,000k contextTools
$1.475 input / 1M$4.425 output / 1M
x-ai

SpaceXAI: Grok Build 0.1

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

Chat256k contextTools
$1 input / 1M$2 output / 1M
google

Google: Gemini 3.5 Flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

Chat1,048.576k contextTools
$1.5 input / 1M$9 output / 1M
google

Google: Gemini 3.5 Flash (batch)

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

Chat1,048.576k contextTools
$0.75 input / 1M$4.5 output / 1M
perceptron

Perceptron: Perceptron Mk1

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...

Chat32.768k context
$0.15 input / 1M$1.5 output / 1M
google

Google: Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Chat1,048.576k contextTools
$0.25 input / 1M$1.5 output / 1M
google

Google: Gemini 3.1 Flash Lite (batch)

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Chat1,048.576k contextTools
$0.125 input / 1M$0.75 output / 1M
openai

OpenAI: GPT Chat Latest

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...

Chat400k contextTools
$5 input / 1M$30 output / 1M
x-ai

SpaceXAI: Grok 4.3

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

Chat1,000k contextTools
$1.25 input / 1M$2.5 output / 1M
x-ai

SpaceXAI: Grok 4.3 (batch)

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

Chat1,000k contextTools
$1 input / 1M$2 output / 1M
mistralai

Mistral: Mistral Medium 3.5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Chat262.144k contextTools
$1.5 input / 1M$7.5 output / 1M
mistralai

Mistral: Mistral Medium 3.5 (batch)

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Chat262.144k contextTools
$0.75 input / 1M$3.75 output / 1M
nvidia

NVIDIA: Nemotron 3 Nano Omni (free)

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

Chat256k contextTools
$0 input / 1M$0 output / 1M

Use one model or an evaluated route.

Automatic mode chooses from your allowed, evaluated models. Catalog inclusion alone does not add a model to that routing pool.

Learn about routing →

Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.