AI models

Every way
to use the major models.

Closed models like Claude and GPT — link to the cheapest API provider. Open-weights like Llama, Kimi, DeepSeek — choose hosted inference or self-host on rented GPUs.

217 tracked · 0 open weights · 217 closed APIs · cheapest input $0.03/M
Quality × Price

Find the sweet spot.

Higher = stronger benchmark composite · further left = cheaper input

Loading...

217 models match — reset filters

Closed / API-only models.

Direct API, aggregator (OpenRouter, Bedrock), or chat UI.

Google: Gemini 3.8 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic task...

Anthropic: Claude Fable 5.1

multimodal
by Anthropic · 1,000,000 ctx

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, a...

?

Ox Alpha

multimodal
by Stealth · 1,048,576 ctx

Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon softwa...

Google: Gemini 3.7 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed f...

SpaceXAI: Grok 4.6

multimodal
by xAI · 500,000 ctx

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

?

Sakana: Sakana Namazu

multimodal
by Sakana · 262,144 ctx

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language...

?

Thinking Machines: Inkling Small

multimodal
by Thinkingmachines · 524,288 ctx

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B to...

Claude Opus 5

multimodal
by Anthropic · 1,000,000 ctx

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at ...

Google: Gemini 3.5 Flash-Lite

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.5 Flash-Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute ...

Google: Gemini 3.6 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to pro...

?

Meituan: LongCat 2.0

text
by Meituan · 1,048,756 ctx

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for co...

?

Poolside: Laguna S 2.1

text
by Poolside · 1,048,576 ctx

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with ...

?

Auto Router (Beta)

multimodal
by Openrouter · 2,000,000 ctx

Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#...

OpenAI: GPT-5.6 Luna

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as ch...

OpenAI: GPT-5.6 Luna Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` ...

OpenAI: GPT-5.6 Sol

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is p...

OpenAI: GPT-5.6 Sol Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set...

OpenAI: GPT-5.6 Terra

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. ...

OpenAI: GPT-5.6 Terra Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mod...

?

Venice: Uncensored

24B
by Cognitivecomputations · 128,000 ctx

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in ...

xAI: Grok 4.5

multimodal
by xAI · 500,000 ctx

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

?

Poolside: Laguna XS 2.1

text
by Poolside · 262,144 ctx

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from thei...

Anthropic: Claude Sonnet 5

multimodal
by Anthropic · 1,000,000 ctx

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It suppo...

Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

multimodal
by Google DeepMind · 65,536 ctx

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity dev...

?

Sakana: Fugu Ultra

multimodal
by Sakana · 1,000,000 ctx

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-age...

?

Poolside: Laguna M.1

text
by Poolside · 262,144 ctx

Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. De...

?

Poolside: Laguna XS.2

text
by Poolside · 262,144 ctx

Laguna XS.2 is the second-generation model in the XS size class from [Poolside](https://poolside.ai/), their efficient coding agent serie...

Google: Nano Banana 2 (Gemini 3.1 Flash Image)

multimodal
by Google DeepMind · 131,072 ctx

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-le...

Google: Nano Banana Pro (Gemini 3 Pro Image)

multimodal
by Google DeepMind · 65,536 ctx

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana ...

Anthropic: Claude Fable 5

multimodal
by Anthropic · 1,000,000 ctx

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file ...

?

OpenRouter: Fusion

text
by Openrouter · 128,000 ctx

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel wit...

Anthropic: Claude Opus 4.8

multimodal
by Anthropic · 1,000,000 ctx

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with t...

AI21: Jamba Large 1.7

text
by Ai21 · 256,000 ctx

Jamba Large 1.7 is the latest model in the Jamba open family, offering improvements in grounding, instruction-following, and overall effi...

Amazon: Nova 2 Lite

multimodal
by Amazon · 1,000,000 ctx

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. ...

Amazon: Nova Lite 1.0

multimodal
by Amazon · 300,000 ctx

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to ...

Amazon: Nova Micro 1.0

text
by Amazon · 128,000 ctx

Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low c...

Amazon: Nova Premier 1.0

multimodal
by Amazon · 1,000,000 ctx

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for dis...

Amazon: Nova Pro 1.0

multimodal
by Amazon · 300,000 ctx

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide ...

Anthropic: Claude 3 Haiku

multimodal
by Anthropic · 200,000 ctx

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. S...

Anthropic: Claude Opus 4

multimodal
by Anthropic · 200,000 ctx

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-runnin...

Anthropic: Claude Opus 4.1

multimodal
by Anthropic · 200,000 ctx

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic task...

Anthropic: Claude Opus 4.5

multimodal
by Anthropic · 200,000 ctx

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon c...

Anthropic: Claude Opus 4.6

multimodal
by Anthropic · 1,000,000 ctx

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire...

Anthropic: Claude Sonnet 4

multimodal
by Anthropic · 1,000,000 ctx

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with...

Anthropic: Claude Sonnet 4.5

multimodal
by Anthropic · 1,000,000 ctx

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers st...

?

Auto Router

multimodal
by Openrouter · 2,000,000 ctx

Your prompt will be processed by a meta-model and routed to one of dozens of models (see below), optimizing for the best possible output....

?

Body Builder (beta)

text
by Openrouter · 128,000 ctx

Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI mod...

Cohere: Command A

text
by Cohere · 256,000 ctx

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, mult...

Cohere: Command R (08-2024)

text
by Cohere · 128,000 ctx

command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmente...

Cohere: Command R+ (08-2024)

text
by Cohere · 128,000 ctx

command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower l...

Cohere: Command R7B (12-2024)

text
by Cohere · 128,000 ctx

Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, an...

?

Free Models Router

multimodal
by Openrouter · 200,000 ctx

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenR...

Google: Gemini 2.0 Flash

multimodal
by Google DeepMind · 1,000,000 ctx

Gemini Flash 2.0 offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while...

Google: Gemini 2.0 Flash Lite

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), ...

Google: Gemini 2.5 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and sci...

Google: Gemini 2.5 Flash Lite

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It ...

Google: Gemini 3.1 Flash Lite

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text,...

Google: Gemini 3.5 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed....

Google: Gemma 2 27B

27B
by Google DeepMind · 8,192 ctx

Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). ...

Google: Gemma 3 12B

12B
by Google DeepMind · 131,072 ctx

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, unders...

Google: Gemma 3n 4B

4B
by Google DeepMind · 32,768 ctx

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It support...

Google: Gemma 4 26B A4B

26B
by Google DeepMind · 262,144 ctx

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B...

Google: Gemma 4 31B

31B
by Google DeepMind · 262,144 ctx

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K ...

Google: Nano Banana (Gemini 2.5 Flash Image)

multimodal
by Google DeepMind · 32,768 ctx

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual...

?

Inflection: Inflection 3 Pi

text
by Inflection · 8,000 ctx

Inflection 3 Pi powers Inflection's [Pi](https://pi.ai) chatbot, including backstory, emotional intelligence, productivity, and safety. I...

?

Inflection: Inflection 3 Productivity

text
by Inflection · 8,000 ctx

Inflection 3 Productivity is optimized for following instructions. It is better for tasks requiring JSON output or precise adherence to p...

?

AionLabs: Aion 3.5

text
by Aion Labs · 262,144 ctx

Aion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative g...

?

AionLabs: Aion 3.5 Mini

text
by Aion Labs · 262,144 ctx

Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, l...

Anthropic: Claude Haiku 5.5

multimodal
by Anthropic · 1,000,000 ctx

Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use....

Anthropic: Claude Opus 5.5

multimodal
by Anthropic · 1,000,000 ctx

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. I...

Anthropic: Claude Sonnet 5.5

multimodal
by Anthropic · 1,000,000 ctx

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is ...

Cohere: Command A+

multimodal
by Cohere · 192,000 ctx

Command A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supp...

?

DeepSeek: DeepSeek Flash Latest

multimodal
by ~deepseek · 1,048,576 ctx

This model always redirects to the latest model in the DeepSeek Flash family.

?

DeepSeek: DeepSeek Pro Latest

text
by ~deepseek · 1,048,576 ctx

This model always redirects to the latest model in the DeepSeek Pro family.

Google: Nano Banana 2.1

multimodal
by Google DeepMind · 65,536 ctx

Nano Banana 2.1 (Gemini Nano Banana 2.1) is Google's image generation and editing model on the Flash tier, succeeding Nano Banana 2 and N...

?

inclusionAI: Ling 3.1 Flash

text
by Inclusionai · 262,144 ctx

Ling 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.

?

Inference.net: Schematron V2 Small

text
by Inference Net · 128,000 ctx

Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex sch...

?

Inference.net: Schematron V2 Turbo

text
by Inference Net · 128,000 ctx

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extract...

Mistral: Mistral Large 4

multimodal
by Mistral AI · 524,288 ctx

Mistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads....

?

Nex AGI: Nex-N2.5-Mini

multimodal
by Nex Agi · 262,144 ctx

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual fee...

?

Nex AGI: Nex-N2.5-Pro

multimodal
by Nex Agi · 262,144 ctx

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual fee...

NVIDIA: Switchyard

text
by Nvidia · 1,000,000 ctx

Switchyard is an open-source model router that switches between multiple models to optimize the cost of requests. By default it will use ...

OpenAI: GPT-6.1 Sol

multimodal
by OpenAI · 1,050,000 ctx

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agent...

OpenAI: GPT-6.1 Sol Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-6.1 Sol Pro is the same underlying model as [GPT-6.1 Sol](https://openrouter.ai/openai/gpt-6.1-sol), served with `reasoning.mode` set...

OpenAI: GPT-6 Astra

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep rese...

OpenAI: GPT-6 Astra Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set...

OpenAI: GPT-6 Luna

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and laten...

OpenAI: GPT-6 Luna Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to...

OpenAI: GPT-6 Sol

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT...

OpenAI: GPT-6 Sol Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `p...

?

OpenAI GPT Astra Latest

multimodal
by ~openai · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT Astra family.

?

OpenAI GPT Luna Latest

multimodal
by ~openai · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT Luna family.

?

OpenAI GPT Sol Latest

multimodal
by ~openai · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT Sol family.

?

OpenAI GPT Terra Latest

multimodal
by ~openai · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT Terra family.

?

Pareto

multimodal
by Unbiased · 262,144 ctx

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance acr...

?

Pareto 26.10 Preview

multimodal
by Unbiased · 1,048,576 ctx

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance acr...

?

Perceptron: Perceptron Mk1.5

multimodal
by Perceptron · 36,864 ctx

Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answer...

?

PrismML: Ternary Bonsai 2 27B

multimodal
by Prism ML · 262,144 ctx

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and...

Qwen: Qwen3.8 Max Prime

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price poin...

Qwen: Qwen3.8 Omni Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audi...

?

Sakana: Fugu Max

multimodal
by Sakana · 1,000,000 ctx

Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent o...

?

Sakana: Fugu Ultra v2

multimodal
by Sakana · 1,000,000 ctx

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-...

?

Space Bunny Alpha

multimodal
by Stealth · 1,000,000 ctx

Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support...

SpaceXAI: Grok 4.7

multimodal
by xAI · 500,000 ctx

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at lo...

?

TypeSafe: Jev Router

multimodal
by Typesafe · 1,000,000 ctx

Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. It runs on [Jev](https://openr...

?

Union Alpha

multimodal
by Stealth · 262,144 ctx

Union Alpha is a multimodal model built for research, coding, and agentic workflows, while delivering frontier-level performance across a...

Upstage: Solar Mini 4

text
by Upstage · 524,288 ctx

Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K...

Xiaomi: MiMo-V2.6-Flash

multimodal
by Xiaomi · 1,048,576 ctx

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parame...

Xiaomi: MiMo-V2.6-Pro

multimodal
by Xiaomi · 1,048,576 ctx

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ce...

Xiaomi: MiMo-V2.6-Pro-UltraSpeed

multimodal
by Xiaomi · 1,048,576 ctx

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6...

Z.ai: GLM 5.3 FlashX

multimodal
by Zhipu AI · 1,048,576 ctx

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 toke...

Z.ai: GLM 5.3 Prime

text
by Zhipu AI · 1,000,000 ctx

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput...

OpenAI: GPT-3.5 Turbo

text
by OpenAI · 16,385 ctx

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and tradition...

OpenAI: GPT-3.5 Turbo (older v0613)

text
by OpenAI · 4,095 ctx

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and tradition...

OpenAI: GPT-3.5 Turbo 16k

text
by OpenAI · 16,385 ctx

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single reque...

OpenAI: GPT-3.5 Turbo Instruct

text
by OpenAI · 4,095 ctx

This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Se...

OpenAI: GPT-4

text
by OpenAI · 8,191 ctx

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy tha...

OpenAI: GPT-4 (older v0314)

text
by OpenAI · 8,191 ctx

GPT-4-0314 is the first version of GPT-4 released, with a context length of 8,192 tokens, and was supported until June 14. Training data:...

OpenAI: GPT-4 Turbo (older v1106)

text
by OpenAI · 128,000 ctx

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to ...

OpenAI: GPT-4.1

multimodal
by OpenAI · 1,047,576 ctx

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-contex...

OpenAI: GPT-4.1 Mini

multimodal
by OpenAI · 1,047,576 ctx

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 ...

OpenAI: GPT-4.1 Nano

multimodal
by OpenAI · 1,047,576 ctx

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performa...

OpenAI: GPT-4o (2024-05-13)

multimodal
by OpenAI · 128,000 ctx

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligen...

OpenAI: GPT-4o (2024-08-06)

multimodal
by OpenAI · 128,000 ctx

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respo...

OpenAI: GPT-4o (2024-11-20)

multimodal
by OpenAI · 128,000 ctx

The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improv...

OpenAI: GPT-4o-mini (2024-07-18)

multimodal
by OpenAI · 128,000 ctx

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. ...

OpenAI: GPT-5 Chat

multimodal
by OpenAI · 128,000 ctx

GPT-5 Chat is designed for advanced, natural, multimodal, and context-aware conversations for enterprise applications.

OpenAI: GPT-5 Codex

multimodal
by OpenAI · 400,000 ctx

GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactiv...

OpenAI: GPT-5 Image

multimodal
by OpenAI · 400,000 ctx

[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It o...

OpenAI: GPT-5 Image Mini

multimodal
by OpenAI · 400,000 ctx

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with...

OpenAI: GPT-5 Mini

multimodal
by OpenAI · 400,000 ctx

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following a...

OpenAI: GPT-5 Nano

multimodal
by OpenAI · 400,000 ctx

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low late...

OpenAI: GPT-5 Pro

multimodal
by OpenAI · 400,000 ctx

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized f...

OpenAI: GPT-5.1

multimodal
by OpenAI · 400,000 ctx

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adheren...

OpenAI: GPT-5.1 Chat

multimodal
by OpenAI · 128,000 ctx

GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong genera...

OpenAI: GPT-5.1-Codex

multimodal
by OpenAI · 400,000 ctx

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both intera...

OpenAI: GPT-5.1-Codex-Max

multimodal
by OpenAI · 400,000 ctx

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is base...

OpenAI: GPT-5.1-Codex-Mini

multimodal
by OpenAI · 400,000 ctx

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

OpenAI: GPT-5.2

multimodal
by OpenAI · 400,000 ctx

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1...

OpenAI: GPT-5.2 Chat

multimodal
by OpenAI · 128,000 ctx

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong gener...

OpenAI: GPT-5.2 Pro

multimodal
by OpenAI · 400,000 ctx

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. I...

OpenAI: GPT-5.2-Codex

multimodal
by OpenAI · 400,000 ctx

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both in...

OpenAI: GPT-5.3 Chat

multimodal
by OpenAI · 128,000 ctx

GPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful...

OpenAI: GPT-5.3-Codex

multimodal
by OpenAI · 400,000 ctx

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex wi...

OpenAI: GPT-5.4

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window ...

OpenAI: GPT-5.4 Image 2

multimodal
by OpenAI · 272,000 ctx

[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabiliti...

OpenAI: GPT-5.4 Mini

multimodal
by OpenAI · 400,000 ctx

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It suppor...

OpenAI: GPT-5.4 Nano

multimodal
by OpenAI · 400,000 ctx

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks...

OpenAI: GPT-5.4 Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex,...

OpenAI: GPT-5.5

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher relia...

OpenAI: GPT-5.5 Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a ...

OpenAI: gpt-oss-safeguard-20b

20B
by OpenAI · 131,072 ctx

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts ...

OpenAI: o1

multimodal
by OpenAI · 200,000 ctx

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is t...

OpenAI: o1-pro

multimodal
by OpenAI · 200,000 ctx

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro mod...

OpenAI: o3

multimodal
by OpenAI · 200,000 ctx

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It ...

OpenAI: o3 Deep Research

multimodal
by OpenAI · 200,000 ctx

o3-deep-research is OpenAI's advanced model for deep research, designed to tackle complex, multi-step research tasks. Note: This model a...

OpenAI: o3 Mini

text
by OpenAI · 200,000 ctx

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

OpenAI: o3 Mini High

text
by OpenAI · 200,000 ctx

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient langua...

OpenAI: o3 Pro

multimodal
by OpenAI · 200,000 ctx

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro mode...

OpenAI: o4 Mini

multimodal
by OpenAI · 200,000 ctx

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multim...

OpenAI: o4 Mini Deep Research

multimodal
by OpenAI · 200,000 ctx

o4-mini-deep-research is OpenAI's faster, more affordable deep research model—ideal for tackling complex, multi-step research tasks. Not...

OpenAI: o4 Mini High

multimodal
by OpenAI · 200,000 ctx

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reason...

?

Owl Alpha

text
by Openrouter · 1,048,756 ctx

Owl Alpha is a high-performance foundation model designed for agentic workloads. Natively supports tool use, and long-context tasks, with...

?

Pareto Code Router

text
by Openrouter · 2,000,000 ctx

The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) c...

Perplexity: Sonar

multimodal
by Perplexity · 127,072 ctx

Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed ...

Perplexity: Sonar Deep Research

text
by Perplexity · 128,000 ctx

Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It aut...

Perplexity: Sonar Pro

multimodal
by Perplexity · 200,000 ctx

Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing...

Perplexity: Sonar Pro Search

multimodal
by Perplexity · 200,000 ctx

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is d...

Perplexity: Sonar Reasoning Pro

multimodal
by Perplexity · 128,000 ctx

Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing...

xAI: Grok 4.20

multimodal
by xAI · 2,000,000 ctx

Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest halluci...

xAI: Grok 4.20 Multi-Agent

multimodal
by xAI · 2,000,000 ctx

Grok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in paral...

xAI: Grok 4.3

multimodal
by xAI · 1,000,000 ctx

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instructi...

xAI: Grok Build 0.1

multimodal
by xAI · 256,000 ctx

Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inp...

Claude Opus 4.7

text
by Anthropic · Claude · 200,000 ctx

Frontier reasoning and long-form coding from Anthropic.

Claude Sonnet 4.6

text
by Anthropic · Claude · 200,000 ctx

Best price-performance from Anthropic. Default for production agents.

Claude Haiku 4.5

text
by Anthropic · Claude · 200,000 ctx

Fast, cheap Claude variant for high-throughput inference.

GPT-5

text
by OpenAI · GPT · 256,000 ctx

OpenAI's frontier multimodal reasoning model.

Gemini 2.5 Pro

multimodal
by Google DeepMind · Gemini · 1,000,000 ctx

Google's frontier reasoning model with native 1M-token context.

Grok 3

multimodal
by xAI · Grok · 1,000,000 ctx

xAI's frontier model with built-in DeepSearch + real-time X integration.

Claude 3.5 Haiku

text
by Anthropic · Claude · 200,000 ctx

Fast/cheap Claude 3.5 variant — production fallback for Haiku 4.5.

GPT-4o Mini

multimodal
by OpenAI · GPT · 128,000 ctx

Cheap multimodal default — replaced GPT-3.5 Turbo for low-cost workloads.

Claude 3.5 Sonnet

text
by Anthropic · Claude · 200,000 ctx

Anthropic's 3.5 generation — still in active production.

Gemini 1.5 Flash

multimodal
by Google DeepMind · Gemini · 1,000,000 ctx

Cheap fast Gemini — production default before 2.0/2.5 Flash.

GPT-4o

multimodal
by OpenAI · GPT · 128,000 ctx

OpenAI's multimodal model — text, vision, audio in one.

GPT-4 Turbo

text
by OpenAI · GPT · 128,000 ctx

OpenAI's pre-GPT-5 flagship — still extensively deployed.

Gemini 1.5 Pro

multimodal
by Google DeepMind · Gemini · 1,000,000 ctx

Google's pre-2.5 frontier — 2M context launched here.

?

Z.ai: GLM Flash Latest

multimodal
by ~z Ai · 1,310,720 ctx

This model always redirects to the latest model in the GLM Flash family.

?

Z.ai: GLM Latest

text
by ~z Ai · 1,048,576 ctx

This model always redirects to the latest GLM model from Z.ai.

?

DeepSeek V4 Flash Latest

text
by ~deepseek · 1,048,576 ctx

This model always redirects to the latest model in the DeepSeek V4 Flash family.

Claude Opus 5 (Fast)

multimodal
by Anthropic · 1,000,000 ctx

Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regul...

?

xAI: Grok Latest

multimodal
by ~x Ai · 500,000 ctx

This model always redirects to the latest Grok model from xAI.

?

Anthropic: Claude Fable Latest

multimodal
by ~anthropic · 1,000,000 ctx

This model always redirects to the latest model in the Claude Fable family.

Anthropic: Claude Opus 4.8 (Fast)

multimodal
by Anthropic · 1,000,000 ctx

Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to r...

Anthropic Claude Haiku Latest

multimodal
by Anthropic · 200,000 ctx

This model always redirects to the latest model in the Anthropic Claude Haiku family.

Anthropic Claude Sonnet Latest

multimodal
by Anthropic · 1,000,000 ctx

This model always redirects to the latest model in the Anthropic Claude Sonnet family.

Anthropic: Claude Opus 4.6 (Fast)

multimodal
by Anthropic · 1,000,000 ctx

Fast-mode variant of [Opus 4.6](/anthropic/claude-opus-4.6) - identical capabilities with higher output speed at premium 6x pricing. Lea...

Anthropic: Claude Opus 4.7 (Fast)

multimodal
by Anthropic · 1,000,000 ctx

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Lea...

Anthropic: Claude Opus Latest

multimodal
by Anthropic · 1,000,000 ctx

This model always redirects to the latest model in the Claude Opus family.

?

Google Gemini Flash Latest

multimodal
by ~google · 1,048,576 ctx

This model always redirects to the latest model in the Google Gemini Flash family.

?

Google Gemini Pro Latest

multimodal
by ~google · 1,048,576 ctx

This model always redirects to the latest model in the Google Gemini Pro family.

Google: Gemini 2.5 Flash Lite Preview 09-2025

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It ...

Google: Gemini 2.5 Pro Preview 05-06

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It emplo...

Google: Gemini 2.5 Pro Preview 06-05

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It emplo...

Google: Gemini 3 Flash Preview

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance....

Google: Gemini 3.1 Flash Lite Preview

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite...

Google: Gemini 3.1 Pro Preview

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic relia...

Google: Gemini 3.1 Pro Preview Custom Tools

multimodal
by Google DeepMind · 1,048,756 ctx

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a gener...

Google: Lyria 3 Clip Preview

multimodal
by Google DeepMind · 1,048,576 ctx

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemin...

Google: Lyria 3 Pro Preview

multimodal
by Google DeepMind · 1,048,576 ctx

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. ...

Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)

multimodal
by Google DeepMind · 131,072 ctx

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, deliverin...

Google: Nano Banana Pro (Gemini 3 Pro Image Preview)

multimodal
by Google DeepMind · 65,536 ctx

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana ...

OpenAI GPT Latest

multimodal
by OpenAI · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT family.

OpenAI GPT Mini Latest

multimodal
by OpenAI · 400,000 ctx

This model always redirects to the latest model in the OpenAI GPT Mini family.

OpenAI: GPT Chat Latest

multimodal
by OpenAI · 400,000 ctx

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. ...

OpenAI: GPT-4 Turbo Preview

text
by OpenAI · 128,000 ctx

The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Traini...

OpenAI: GPT-4o Search Preview

text
by OpenAI · 128,000 ctx

GPT-4o Search Previewis a specialized model for web search in Chat Completions. It is trained to understand and execute web search queries.

OpenAI: GPT-4o-mini Search Preview

text
by OpenAI · 128,000 ctx

GPT-4o mini Search Preview is a specialized model for web search in Chat Completions. It is trained to understand and execute web search ...