THE MODEL CATALOG

Chat model APIs & token pricing

Compare text chat models, input and output token prices, context windows and tool support. Use one API key with the OpenAI-compatible chat endpoint.

435 models · Page 4 of 10

~anthropic

Anthropic: Claude Haiku Latest

This model always redirects to the latest model in the Claude Haiku family.

Chat200k contextTools
$1 input / 1M$5 output / 1M
~openai

OpenAI: GPT Mini Latest

This model always redirects to the latest model in the GPT Mini family.

Chat400k contextTools
$0.75 input / 1M$4.5 output / 1M
~google

Google: Gemini Pro Latest

This model always redirects to the latest model in the Gemini Pro family.

Chat1,048.576k contextTools
$2 input / 1M$12 output / 1M
~moonshotai

MoonshotAI: Kimi Latest

This model always redirects to the latest model in the Kimi family.

Chat1,048.576k contextTools
$1.5 input / 1M$7.5 output / 1M
~google

Google: Gemini Flash Latest

This model always redirects to the latest model in the Gemini Flash family.

Chat1,048.576k contextTools
$0.75 input / 1M$3.75 output / 1M
~anthropic

Anthropic: Claude Sonnet Latest

This model always redirects to the latest model in the Claude Sonnet family.

Chat1,000k contextTools
$2 input / 1M$10 output / 1M
qwen

Qwen: Qwen3.5 Plus 2026-04-20

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...

Chat1,000k contextTools
$0.3 input / 1M$1.8 output / 1M
qwen

Qwen: Qwen3.6 Flash

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...

Chat1,000k contextTools
$0.1875 input / 1M$1.125 output / 1M
qwen

Qwen: Qwen3.6 35B A3B

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

Chat262.144k contextTools
$0.15 input / 1M$1 output / 1M
qwen

Qwen: Qwen3.6 Max Preview

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

Chat262.144k contextTools
$1.027 input / 1M$6.162 output / 1M
qwen

Qwen: Qwen3.6 27B

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

Chat262.144k contextTools
$0.3 input / 1M$2 output / 1M
openai

OpenAI: GPT-5.5 Pro

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

Chat1,050k contextTools
$30 input / 1M$180 output / 1M
openai

OpenAI: GPT-5.5 Pro (batch)

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

Chat1,050k contextTools
$15 input / 1M$90 output / 1M
openai

OpenAI: GPT-5.5

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Chat1,050k contextTools
$5 input / 1M$30 output / 1M
openai

OpenAI: GPT-5.5 (batch)

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Chat1,050k contextTools
$2.5 input / 1M$15 output / 1M
deepseek

DeepSeek: DeepSeek V4 Pro 0423

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Chat1,048.576k contextTools
$0.95526 input / 1M$1.91052 output / 1M
deepseek

DeepSeek: DeepSeek V4 Flash 0423

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

Chat1,048.576k contextTools
$0.088606 input / 1M$0.177212 output / 1M
tencent

Tencent: Hy3 preview

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...

Chat262.144k contextTools
$0.18 input / 1M$0.6 output / 1M
xiaomi

Xiaomi: MiMo-V2.5-Pro

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....

Chat1,050k contextTools
$0.435 input / 1M$0.87 output / 1M
xiaomi

Xiaomi: MiMo-V2.5

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

Chat1,050k contextTools
$0.14 input / 1M$0.28 output / 1M
~anthropic

Anthropic: Claude Opus Latest

This model always redirects to the latest model in the Claude Opus family.

Chat1,000k contextTools
$5 input / 1M$25 output / 1M
openrouter

Pareto Code Router

The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles. Set min_coding_score between 0 and 1 on the [pareto-router plugin](https://openrouter.ai/docs/guides/routing/routers/pareto-router#the-min_coding_score-parameter) to control how...

Chat2,000k context
Price depends on the selected model
moonshotai

MoonshotAI: Kimi K2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

Chat262.144k contextTools
$0.95 input / 1M$4 output / 1M
anthropic

Anthropic: Claude Opus 4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

Chat1,000k contextTools
$5 input / 1M$25 output / 1M
anthropic

Anthropic: Claude Opus 4.7 (batch)

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

Chat1,000k contextTools
$2.5 input / 1M$12.5 output / 1M
z-ai

Z.ai: GLM 5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

Chat204.8k contextTools
$0.966 input / 1M$3.036 output / 1M
google

Google: Gemma 4 26B A4B

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Chat262.144k contextTools
$0.09 input / 1M$0.3 output / 1M
google

Google: Gemma 4 26B A4B (free)

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
google

Google: Gemma 4 31B

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Chat262.144k contextTools
$0.09 input / 1M$0.34 output / 1M
google

Google: Gemma 4 31B (free)

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M
qwen

Qwen: Qwen3.6 Plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

Chat1,000k contextTools
$0.325 input / 1M$1.95 output / 1M
z-ai

Z.ai: GLM 5V Turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

Chat202.752k contextTools
$1.2 input / 1M$4 output / 1M
arcee-ai

Arcee AI: Trinity Large Thinking

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...

Chat262.144k contextTools
$0.25 input / 1M$0.8 output / 1M
x-ai

SpaceXAI: Grok 4.20 Multi-Agent

Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

Chat2,000k context
$1.25 input / 1M$2.5 output / 1M
x-ai

SpaceXAI: Grok 4.20

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

Chat2,000k contextTools
$1.25 input / 1M$2.5 output / 1M
google

Google: Lyria 3 Pro Preview

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...

ChatAudio1,048.576k context
$0 input / 1M$0 output / 1M
google

Google: Lyria 3 Clip Preview

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...

ChatAudio1,048.576k context
$0 input / 1M$0 output / 1M
rekaai

Reka Edge

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,...

Chat16.384k contextTools
$0.1 input / 1M$0.1 output / 1M
minimax

MiniMax: MiniMax M2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

Chat204.8k contextTools
$0.3 input / 1M$1.2 output / 1M
openai

OpenAI: GPT-5.4 Nano

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

Chat400k contextTools
$0.2 input / 1M$1.25 output / 1M
openai

OpenAI: GPT-5.4 Nano (batch)

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

Chat400k contextTools
$0.1 input / 1M$0.625 output / 1M
openai

OpenAI: GPT-5.4 Mini

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

Chat400k contextTools
$0.75 input / 1M$4.5 output / 1M
openai

OpenAI: GPT-5.4 Mini (batch)

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

Chat400k contextTools
$0.375 input / 1M$2.25 output / 1M
mistralai

Mistral: Mistral Small 4

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

Chat262.144k contextTools
$0.15 input / 1M$0.6 output / 1M
mistralai

Mistral: Mistral Small 4 (batch)

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

Chat262.144k contextTools
$0.075 input / 1M$0.3 output / 1M
z-ai

Z.ai: GLM 5 Turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...

Chat202.752k contextTools
$1.2 input / 1M$4 output / 1M
nvidia

NVIDIA: Nemotron 3 Super

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Chat262.144k contextTools
$0.08 input / 1M$0.45 output / 1M
nvidia

NVIDIA: Nemotron 3 Super (free)

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Chat262.144k contextTools
$0 input / 1M$0 output / 1M

Use one model or an evaluated route.

Automatic mode chooses from your allowed, evaluated models. Catalog inclusion alone does not add a model to that routing pool.

Learn about routing →

Catalog prices refresh hourly. This public view may lag the latest sync by up to five minutes. Capabilities are provider-declared; availability can change.