THE MODEL CATALOG

Chat model APIs & token pricing

Compare text chat models, input and output token prices, context windows and tool support. Use one API key with the OpenAI-compatible chat endpoint.

435 models · Page 2 of 10

z-ai

Z.ai: GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Chat1,310.72k contextTools
$0.84 input / 1M$2.64 output / 1M
z-ai

Z.ai: GLM 5.3 (batch)

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Chat1,048.576k contextTools
$0.7 input / 1M$2.2 output / 1M
qwen

Qwen: Qwen3.8 27B

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

Chat1,000k contextTools
$0.42 input / 1M$3 output / 1M
qwen

Qwen: Qwen3.8 27B (free)

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
dots-studio

Dots Studio: Dots3-Note Preview (free)

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is...

Chat512k contextTools
$0 input / 1M$0 output / 1M
google

Google: Gemini 3.7 Flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Chat1,048.576k contextTools
$0.75 input / 1M$3.75 output / 1M
google

Google: Gemini 3.7 Flash (batch)

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Chat1,048.576k contextTools
$0.375 input / 1M$1.875 output / 1M
bytedance-seed

ByteDance Seed: Seed 2.1 Turbo

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...

Chat262.144k contextTools
$0.5 input / 1M$2.5 output / 1M
qwen

Qwen: Qwen3.8 2.4T A95B

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

Chat1,048.576k contextTools
$2 input / 1M$6 output / 1M
bytedance-seed

ByteDance Seed: Seed-2.0-Code

Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude...

Chat262.144k contextTools
$0.5 input / 1M$3 output / 1M
deepseek

DeepSeek: DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Chat1,048.576k contextTools
$0.624 input / 1M$2.88 output / 1M
deepseek

DeepSeek: DeepSeek V4 Pro 0813 (batch)

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Chat1,048.576k contextTools
$0.66 input / 1M$1.98 output / 1M
x-ai

SpaceXAI: Grok 4.6

Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7).

Chat500k contextTools
$2 input / 1M$6 output / 1M
liquid

LiquidAI: LFM2.5-2.6B (free)

LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or...

Chat65.536k contextTools
$0 input / 1M$0 output / 1M
nvidia

NVIDIA: Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Chat262.144k contextTools
$0.07 input / 1M$0.2 output / 1M
nvidia

NVIDIA: Nemotron 3.5 Lightning (free)

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Chat1,000k contextTools
$0 input / 1M$0 output / 1M
sakana

Sakana: Sakana Namazu

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...

Chat262.144k contextTools
$0.95 input / 1M$4 output / 1M
upstage

Upstage: Solar Pro 4

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...

Chat524.288k contextTools
$0.09 input / 1M$0.36 output / 1M
meta

Meta: Muse Glimmer 30B

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...

Chat131.072k contextTools
$0.3 input / 1M$1.2 output / 1M
meta

Meta: Muse Glimmer 30B (batch)

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...

Chat131.072k contextTools
$0.175 input / 1M$0.75 output / 1M
meta

Meta: Muse Spark 1.2

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...

Chat1,048.576k contextTools
$1.25 input / 1M$4.25 output / 1M
~deepseek

DeepSeek: DeepSeek V4 Flash Latest

This model always redirects to the latest model in the DeepSeek V4 Flash family.

Chat1,310.72k contextTools
$0.03 input / 1M$1 output / 1M
deepseek

DeepSeek: DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

Chat1,310.72k contextTools
$0.04 input / 1M$0.64 output / 1M
deepseek

DeepSeek: DeepSeek V4 Flash 0731 (batch)

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

Chat1,048.576k contextTools
$0.11 input / 1M$0.33 output / 1M
thinkingmachines

Thinking Machines: Inkling Small

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

Chat1,048.576k contextTools
$0.45 input / 1M$1.2 output / 1M
thinkingmachines

Thinking Machines: Inkling Small (free)

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

Chat1,048.576k contextTools
$0 input / 1M$0 output / 1M
qwen

Qwen: Qwen3.7 Flash

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

Chat1,000k contextTools
$0.03 input / 1M$0.13 output / 1M
anthropic

Anthropic: Claude Opus 5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Chat1,000k contextTools
$5 input / 1M$25 output / 1M
anthropic

Anthropic: Claude Opus 5 (batch)

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Chat1,000k contextTools
$2.5 input / 1M$12.5 output / 1M
inclusionai

inclusionAI: Ling 3.0 Flash

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

Chat262.144k contextTools
$0.021 input / 1M$0.063 output / 1M
poolside

Poolside: Laguna S 2.1

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Chat1,048.576k contextTools
$0.09 input / 1M$0.18 output / 1M
poolside

Poolside: Laguna S 2.1 (free)

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
google

Google: Gemini 3.6 Flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Chat1,048.576k contextTools
$0.75 input / 1M$3.75 output / 1M
google

Google: Gemini 3.6 Flash (batch)

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Chat1,048.576k contextTools
$0.375 input / 1M$1.875 output / 1M
google

Google: Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Chat1,048.576k contextTools
$0.3 input / 1M$2.5 output / 1M
google

Google: Gemini 3.5 Flash Lite (batch)

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Chat1,048.576k contextTools
$0.15 input / 1M$1.25 output / 1M
meituan

Meituan: LongCat 2.0

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...

Chat1,048.756k contextTools
$0.3 input / 1M$1.2 output / 1M
thinkingmachines

Thinking Machines: Inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Chat1,048.576k contextTools
$1 input / 1M$4.05 output / 1M
thinkingmachines

Thinking Machines: Inkling (free)

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Chat1,048.576k contextTools
$0 input / 1M$0 output / 1M
openrouter

Auto Router (Beta)

The experimental version of our Auto Router where we test new improvements. Use it to get the latest and greatest version of our general purpose auto router, but expect beta...

ChatImage2,000k contextTools
Price depends on the selected model
moonshotai

MoonshotAI: Kimi K3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Chat1,048.576k contextTools
$3 input / 1M$15 output / 1M
meta

Meta: Muse Spark 1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

Chat1,048.576k contextTools
$1.25 input / 1M$4.25 output / 1M
kwaipilot

Kwaipilot: KAT-Coder-Pro V2.5

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

Chat262.144k contextTools
$0.74 input / 1M$2.96 output / 1M
openai

OpenAI: GPT-5.6 Luna Pro

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Chat1,050k contextTools
$0.2 input / 1M$1.2 output / 1M
openai

OpenAI: GPT-5.6 Luna Pro (batch)

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Chat1,050k contextTools
$0.1 input / 1M$0.6 output / 1M
openai

OpenAI: GPT-5.6 Luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Chat1,050k contextTools
$0.2 input / 1M$1.2 output / 1M
openai

OpenAI: GPT-5.6 Luna (batch)

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Chat1,050k contextTools
$0.1 input / 1M$0.6 output / 1M
openai

OpenAI: GPT-5.6 Terra Pro

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Chat1,050k contextTools
$2 input / 1M$12 output / 1M

Use one model or an evaluated route.

Automatic mode chooses from your allowed, evaluated models. Catalog inclusion alone does not add a model to that routing pool.

Learn about routing →

Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.