AI models

Every way
to use the major models.

Closed models like Claude and GPT — link to the cheapest API provider. Open-weights like Llama, Kimi, DeepSeek — choose hosted inference or self-host on rented GPUs.

724 tracked · 507 open weights · 217 closed APIs · cheapest input $0.01/M
Quality × Price

Find the sweet spot.

Higher = stronger benchmark composite · further left = cheaper input

Loading...

724 models match

Open-weights models.

Run yourself on cheap GPUs, or use a hosted-inference provider.

Meta: Muse Spark 1.3

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to k...

?

GLM 5.3 Flash

multimodal
by zai-org · 1,048,576 ctx

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. It incorpo...

?

GLM-5.3

text
by zai-org · 1,048,576 ctx

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex co...

Qwen: Qwen3.8 Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, ...

Z.ai: GLM 5.3 Flash

multimodal
by Zhipu AI · 1,310,720 ctx

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse a...

granite-4.2-30b

30B
by IBM Research · 131,072 ctx

Granite-4.2-30B is the flagship reasoning model in the Granite 4.2 family. It delivers the strongest performance across reasoning-intensi...

granite-4.2-3b

3B
by IBM Research · 131,072 ctx

Granite-4.2-3B is the compact reasoning model in the Granite 4.2 family. Despite its small parameter count, it delivers strong performanc...

granite-4.2-8b

9B
by IBM Research · 131,072 ctx

Granite-4.2-8B is the mid-size reasoning model in the Granite 4.2 family. It delivers strong performance on reasoning-intensive tasks by ...

Tencent: Hy-MT2-7B

7B
by Tencent · 8,192 ctx

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language ...

Mistral: Ministral 8B

8B
by Mistral AI · 128,000 ctx

Ministral 8B is an 8B parameter model featuring a unique interleaved sliding-window attention pattern for faster, memory-efficient infere...

Tencent: Hy-MT2-1.8B

8B
by Tencent · 8,192 ctx

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minori...

Tencent: Hy-MT2-30B-A3B

30B
by Tencent · 8,192 ctx

Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and m...

glm-5.3

text
by Zhipu AI · 1,048,576 ctx

Qwen: Qwen3.8 27B

27B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimoda...

Qwen3.8-2.4T-A95B

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.8-2.4T-A95B is Alibaba's most capable Qwen model to date, a 2.4T-parameter sparse MoE with ~95B active parameters. It is built for ...

?

ByteDance Seed: Seed 2.1 Turbo

multimodal
by Bytedance Seed · 262,144 ctx

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end softwar...

DeepSeek: DeepSeek V4 Pro 0813

text
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Qwen3.8-2.4T-A95B

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.8-2.4T-A95B is Alibaba’s most capable Qwen model to date, a 2.4T-parameter sparse MoE with ~95B active parameters. It is built for ...

Qwen: Qwen3.8 2.4T A95B

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-...

NVIDIA-Nemotron-3.5-Lightning

text
by Nvidia · 28,672 ctx

NVIDIA Nemotron 3.5 Lightning is NVIDIA's fastest open model for always-on agents and high-volume specialized tasks. It delivers a substa...

NVIDIA: Nemotron 3.5 Lightning

text
by Nvidia · 262,144 ctx

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited f...

Nemotron Lightning 3.5 30B A3B

30B
by Nvidia · 262,144 ctx

Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on...

Meta: Muse Glimmer 30B

30B
by Meta AI · 131,072 ctx

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for a...

Upstage: Solar Pro 4

text
by Upstage · 524,288 ctx

Solar Pro 4 is a large language model from Upstage. It is suited for agentic workflows, office productivity, document-intensive work, and...

?

Ling-3.0-flash

text
by Inclusionai · 131,072 ctx

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The mode...

Meta: Muse Spark 1.2

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documen...

Qwen: Qwen3.8 Max

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a mul...

DeepSeek: DeepSeek V4 Flash 0731

text
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-tra...

Qwen: Qwen3.7 Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer ...

Meta: Muse Spark 1.1

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documen...

?

MoonshotAI: Kimi K3

multimodal
by Moonshot AI · 1,048,576 ctx

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and...

?

Ternary Bonsai 27B

27B
by Prism ML · 262,144 ctx

Kwaipilot: KAT-Coder-Air V2.5

text
by Kwaipilot · 256,000 ctx

Kwaipilot: KAT-Coder-Pro V2.5

text
by Kwaipilot · 256,000 ctx

Gemma 4 12B It

12B
by Google DeepMind · 262,144 ctx
?

LFM2.5-8B-A1B

8B
by LiquidAI · 128,000 ctx
?

AionLabs: Aion-3.0

text
by Aion Labs · 131,072 ctx

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative g...

?

AionLabs: Aion-3.0-Mini

text
by Aion Labs · 131,072 ctx

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a colla...

?

Nex AGI: Nex-N2-Mini

multimodal
by Nex Agi · 262,144 ctx

Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text an...

Tencent: Hy3

text
by Tencent · 262,144 ctx

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic w...

?

Ornith-1.0-35B

35B
by Deepreinforce Ai · 262,144 ctx

Ornith-1.0-35B is DeepReinforce's open (MIT-licensed) agentic-coding model: an RL post-train of Qwen3.5-35B-A3B, a 35B-total / ~3B-active...

?

Nex AGI: Nex-N2-Pro

multimodal
by Nex Agi · 262,144 ctx

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 archit...

?

GLM 5.2

text
by zai-org · 1,048,576 ctx

GLM-5.2 introduces a robust 1M-token context and advanced, multi-effort coding capabilities to significantly enhance performance on long-...

Z.ai: GLM 5.2

text
by Zhipu AI · 1,048,576 ctx

GLM-5.2 is Z.ai’s flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-le...

?

Kimi 2.7 Code

multimodal
by Moonshot AI · 262,144 ctx

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding ta...

?

MoonshotAI: Kimi K2.7 Code

multimodal
by Moonshot AI · 262,144 ctx

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reli...

NVIDIA-Nemotron-3-Ultra-550B-A55B

550B
by Nvidia · 262,144 ctx

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It deli...

NVIDIA: Nemotron 3 Ultra

550B
by Nvidia · 1,000,000 ctx

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (...

Nemotron-Content-Safety-3.5

multimodal
by Nvidia · 131,072 ctx

Nemotron Content Safety 3.5 is a multimodal safety classifier developed by NVIDIA. A compact safety model that handles text, images, and...

Qwen: Qwen3.7 Plus

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the se...

MiniMax: MiniMax M3

multimodal
by MiniMax · 1,048,576 ctx

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context ...

StepFun: Step 3.7 Flash

multimodal
by Stepfun · 256,000 ctx

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with ...

?

AionLabs: Aion-1.0

text
by Aion Labs · 131,072 ctx

Aion-1.0 is a multi-model system designed for high performance across various tasks, including reasoning and coding. It is built on DeepS...

?

AionLabs: Aion-1.0-Mini

text
by Aion Labs · 131,072 ctx

Aion-1.0-Mini 32B parameter model is a distilled version of the DeepSeek-R1 model, designed for strong performance in reasoning domains s...

?

AionLabs: Aion-2.0

text
by Aion Labs · 131,072 ctx

Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing ten...

?

AionLabs: Aion-RP 1.0 (8B)

8B
by Aion Labs · 32,768 ctx

Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant ...

?

AlfredPros: CodeLLaMa 7B Instruct Solidity

7B
by Alfredpros · 4,096 ctx

A finetuned 7 billion parameters Code LLaMA - Instruct model to generate Solidity smart contract using 4-bit QLoRA finetuning provided by...

AllenAI: Olmo 3 32B Think

32B
by Allen Institute for AI (AI2) · 65,536 ctx

Olmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instruc...

Arcee AI: Coder Large

text
by Arcee Ai · 32,768 ctx

Coder‑Large is a 32 B‑parameter offspring of Qwen 2.5‑Instruct that has been further trained on permissively‑licensed GitHub, CodeSearchN...

Arcee AI: Maestro Reasoning

text
by Arcee Ai · 131,072 ctx

Maestro Reasoning is Arcee's flagship analysis model: a 32 B‑parameter derivative of Qwen 2.5‑32 B tuned with DPO and chain‑of‑thought RL...

Arcee AI: Spotlight

1B
by Arcee Ai · 131,072 ctx

Spotlight is a 7‑billion‑parameter vision‑language model derived from Qwen 2.5‑VL and fine‑tuned by Arcee AI for tight image‑text groundi...

Arcee AI: Trinity Large Thinking

399B
by Arcee Ai · 262,144 ctx

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, ag...

Arcee AI: Trinity Mini

text
by Arcee Ai · 131,072 ctx

Trinity Mini is a 26B-parameter (3B active) sparse mixture-of-experts language model featuring 128 experts with 8 active per token. Engin...

Arcee AI: Virtuoso Large

text
by Arcee Ai · 131,072 ctx

Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters, tuned to tackle cross‑domain reasoning, creative writing and e...

Arize AI Qwen 2 1.5B Instruct

5B
by Togethercomputer · 32,768 ctx

Baidu: ERNIE 4.5 21B A3B

21B
by Baidu · 131,072 ctx

A sophisticated text-based Mixture-of-Experts (MoE) model featuring 21B total parameters with 3B activated per token, delivering exceptio...

Baidu: ERNIE 4.5 21B A3B Thinking

21B
by Baidu · 131,072 ctx

ERNIE-4.5-21B-A3B-Thinking is Baidu's upgraded lightweight MoE model, refined to boost reasoning depth and quality for top-tier performan...

Baidu: ERNIE 4.5 300B A47B

300B
by Baidu · 131,072 ctx

ERNIE-4.5-300B-A47B is a 300B parameter Mixture-of-Experts (MoE) language model developed by Baidu as part of the ERNIE 4.5 series. It ac...

Baidu: ERNIE 4.5 VL 28B A3B

28B
by Baidu · 131,072 ctx

A powerful multimodal Mixture-of-Experts chat model featuring 28B total parameters with 3B activated per token, delivering exceptional te...

Baidu: ERNIE 4.5 VL 424B A47B

424B
by Baidu · 131,072 ctx

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with...

Baidu: Qianfan-OCR-Fast

multimodal
by Baidu · 65,536 ctx

Qianfan-OCR-Fast is a domain-specific multimodal large model purpose-built for OCR. By leveraging specialized OCR training data while pre...

?

ByteDance Seed: Seed 1.6

200B
by Bytedance Seed · 262,144 ctx

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinki...

?

ByteDance Seed: Seed 1.6 Flash

multimodal
by Bytedance Seed · 262,144 ctx

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It featu...

?

ByteDance Seed: Seed-2.0-Lite

32B
by Bytedance Seed · 262,144 ctx

Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering n...

?

ByteDance Seed: Seed-2.0-Mini

multimodal
by Bytedance Seed · 262,144 ctx

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference...

?

ByteDance: UI-TARS 7B

7B
by Bytedance · 128,000 ctx

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobil...

Deep Cogito: Cogito v2.1 671B

671B
by Deepcogito · 128,000 ctx

Cogito v2.1 671B MoE represents one of the strongest open models globally, matching performance of frontier closed and open models. This ...

DeepSeek-V3-0324

671B
by DeepSeek · 163,840 ctx

DeepSeek-V3-0324, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token, an impro...

DeepSeek-V3.1

text
by DeepSeek · 163,840 ctx

DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase l...

DeepSeek: DeepSeek V3

689B
by DeepSeek · 163,840 ctx

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous vers...

DeepSeek: DeepSeek V3 0324

689B
by DeepSeek · 163,840 ctx

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team...

DeepSeek: DeepSeek V3.1

689B
by DeepSeek · 163,840 ctx

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prom...

DeepSeek: DeepSeek V3.1 Terminus

689B
by DeepSeek · 163,840 ctx

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities whi...

DeepSeek: DeepSeek V3.2

689B
by DeepSeek · 131,072 ctx

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use pe...

DeepSeek: DeepSeek V3.2 Speciale

text
by DeepSeek · 163,840 ctx

DeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance. It builds on D...

DeepSeek: DeepSeek V4 Flash

292B
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated paramete...

DeepSeek: DeepSeek V4 Pro

1602B
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporti...

DeepSeek: R1 0528

671B
by DeepSeek · 163,840 ctx

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced an...

Deepseek Coder 33B Instruct

33B
by DeepSeek · 16,384 ctx

Deepseek OCR 2

text
by DeepSeek · 8,192 ctx

Devstral Small 2505

text
by Mistral AI · 131,072 ctx

EssentialAI Rnj-1 Instruct

text
by Essentialai · 32,768 ctx

EssentialAI: Rnj 1 Instruct

text
by Essentialai · 32,768 ctx

Rnj-1 is an 8B-parameter, dense, open-weight model family developed by Essential AI and trained from scratch with a focus on programming,...

Facebook CWM

text
by Meta AI · 131,072 ctx
?

GLM 5.1

text
by zai-org · 202,752 ctx

GLM-5.1 is Z.ai's next-generation flagship model built for agentic engineering, with stronger coding capabilities and sustained performan...

GLM OCR

text
by Zhipu AI · 131,072 ctx

GLM-4.5-Flash

text
by Zhipu AI · GLM

Lowest-latency, lowest-cost variant of GLM-4.5 on z.ai.

GLM-4.6

355B
by Zhipu AI · GLM · 202,752 ctx

Incremental upgrade on GLM-4.5 — improved reasoning, same context window.

GLM-4.7

9B
by Zhipu AI · GLM · 202,752 ctx

Mid-generation GLM 4.7 released between GLM-4.6 and GLM-5.

GLM-5

355B
by Zhipu AI · GLM · 202,752 ctx

Zhipu's GLM 5 generation — closed flagship between GLM-4.7 and GLM-5.1.

GLM-5 Turbo

text
by Zhipu AI · GLM · 202,752 ctx

Faster, cheaper sibling of GLM-5 on z.ai.

GLM-5.1

32B
by Zhipu AI · GLM · 202,752 ctx

Zhipu's GLM 5.1 series — successor to GLM-5 on z.ai's API.

Gemma 2 9B It

9B
by Google DeepMind · 8,192 ctx

Gemma 2B It

2B
by Google DeepMind · 8,192 ctx

Gemma 3 1B Pt

1B
by Google DeepMind · 32,768 ctx

Gemma 3 1b it

1B
by Google DeepMind · 32,768 ctx

Gemma 3 270M It

0B
by Google DeepMind · 32,768 ctx

Gemma 3 27B It

27B
by Google DeepMind · 65,536 ctx

Gemma 3 27B Pt

27B
by Google DeepMind

Gemma 3 4b it

4B
by Google DeepMind · 65,536 ctx

Gemma 4 E2B-it

text
by Google DeepMind · 131,072 ctx

Gemma 4 E4B-it

text
by Google DeepMind · 131,072 ctx
?

Hermes-3-Llama-3.1-70B

70B
by Nous Research · 131,072 ctx

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better rolepl...

?

Holo3 35B A3b

35B
by Hcompany · 262,144 ctx

IBM: Granite 4.0 Micro

3B
by IBM Research · 131,000 ctx

Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by ...

IBM: Granite 4.1 8B

8B
by IBM Research · 131,072 ctx

Granite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-t...

?

Inception: Mercury 2

text
by Inception · 128,000 ctx

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Me...

?

Kimi K2.5

multimodal
by Moonshot AI · 262,144 ctx

Kimi K2.5 is Moonshot AI's flagship agentic model and a new SOTA open model. It unifies vision and text, thinking and non-thinking modes,...

?

Kimi K2.6

multimodal
by Moonshot AI · 262,144 ctx

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven d...

Kwaipilot: KAT-Coder-Pro V2

text
by Kwaipilot · 256,000 ctx

KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engin...

?

L3.1-70B-Euryale-v2.2

70B
by Sao10k · 131,072 ctx

Euryale 3.1 - 70B v2.2 is a model focused on creative roleplay from Sao10k

LFM2-24B-A2B

24B
by Togethercomputer · 32,768 ctx

LiquidAI: LFM2-24B-A2B

24B
by Liquid · 128,000 ctx

LFM2-24B-A2B is the largest model in the LFM2 family of hybrid architectures designed for efficient on-device deployment. Built as a 24B ...

Llama 3.1 70B

70B
by Meta AI · 131,072 ctx

Llama 3.1 Nemotron 70B Instruct HF

70B
by Nvidia · 32,768 ctx

Llama 4 Scout (17Bx16E)

17B
by Meta AI · 262,144 ctx

Llama 4 Scout Instruct (17Bx16E)

17B
by Meta AI · 1,048,576 ctx

Llama Guard 3 8B

8B
by Meta AI · 131,072 ctx

Llama Guard 3 is a Llama-3.1-8B pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be u...

Llama-3.2-11B-Vision-Instruct

11B
by Meta AI · 131,072 ctx

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It exc...

Magistral Small 2506

text
by Mistral AI · 40,960 ctx
?

Magnum v4 72B

72B
by Anthracite Org · 32,768 ctx

This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anth...

?

Mancer: Weaver (alpha)

text
by Mancer · 8,000 ctx

An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrativ...

Medgemma 27B Text It

27B
by Google DeepMind · 131,072 ctx

DeepSeek: DeepSeek V4.1 Flash

multimodal
by DeepSeek · 1,048,576 ctx

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED)...

DeepSeek V4.1 Flash

multimodal
by DeepSeek · 1,048,576 ctx

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text at up ...

?

Ember-1

multimodal
by fireworks · 1,048,576 ctx

Ember-1 is a specialized model from Fireworks. Built on Kimi K3, it produces shorter reasoning traces, using approximately 40% fewer toke...

?

Inception: Mercury 2.5

text
by Inception · 260,000 ctx

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, ...

?

Ling 3.0 Flash Fin

text
by Inclusionai · 262,144 ctx

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters ou...

?

Ling-3.0-flash-VL

125B
by Inclusionai · 131,072 ctx

The multimodal version built on Ling-3.0-flash — 124B total / ~5.5B active per token, with native text, image, and video understanding. I...

Meta Llama 3 70B Instruct Turbo

70B
by Meta AI · 8,192 ctx

NVIDIA: Nemotron 3.5 Content Safety

multimodal
by Nvidia · 131,072 ctx

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. I...

Qwen: Qwen3.8 Max (0902)

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model ...

?

Tev1 4B Experimental

text
by Together AI · 32,768 ctx

Meta Llama 3 8B Instruct

8B
by Meta AI · 8,192 ctx

Meta Llama 3 8B Instruct Lite

8B
by Meta AI · 8,192 ctx

Meta Llama 3 8B Instruct Reference

8B
by Meta AI · 8,192 ctx

Meta Llama 3.1 405B Instruct

405B
by Meta AI · 4,096 ctx

Meta Llama 3.1 70B Instruct Turbo

70B
by Meta AI · 131,072 ctx

Meta Llama 3.1 8B

8B
by Meta AI · 16,384 ctx

Meta Llama 3.1 8B Instruct Turbo

8B
by Meta AI · 131,072 ctx

Meta Llama 3.2 1B Instruct

1B
by Meta AI · 131,072 ctx

Meta Llama 3.2 3B Instruct

3B
by Meta AI · 131,072 ctx

Meta Llama 3.3 70B Instruct Turbo

70B
by Meta AI · 131,072 ctx

Meta-Llama-3.1-70B-Instruct

70B
by Meta AI · 131,072 ctx

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned g...

Meta-Llama-3.1-8B-Instruct

8B
by Meta AI · 131,072 ctx

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned g...

Meta: Llama 3 70B Instruct

70B
by Meta AI · 8,192 ctx

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 70B instruct-tuned version was optimized for high...

Meta: Llama 3 8B Instruct

8B
by Meta AI · 8,192 ctx

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 8B instruct-tuned version was optimized for high ...

Meta: Llama 4 Maverick

402B
by Meta AI · 1,048,576 ctx

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architec...

Meta: Llama 4 Scout

109B
by Meta AI · 10,000,000 ctx

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of ...

Meta: Llama Guard 4 12B

12B
by Meta AI · 163,840 ctx

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous v...

Microsoft: Phi 4 Mini Instruct

4B
by Microsoft · 131,072 ctx

Phi-4-mini-instruct is a lightweight open model built upon synthetic data and filtered publicly available websites - with a focus on high...

MiniMax M2.7

456B
by MiniMax · 196,608 ctx

Mixture-of-Experts language model. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks...

MiniMax-M2.5

text
by MiniMax · 196,608 ctx

MiniMax M2.5 is built for state-of-the-art coding, agentic tool use, search, and office work, extensively trained with reinforcement lear...

MiniMax: MiniMax M1

456B
by MiniMax · 1,000,000 ctx

MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybr...

MiniMax: MiniMax M2

456B
by MiniMax · 204,800 ctx

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion acti...

MiniMax: MiniMax M2-her

text
by MiniMax · 65,536 ctx

MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn co...

MiniMax: MiniMax M2.1

230B
by MiniMax · 204,800 ctx

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application deve...

MiniMax: MiniMax M2.5

230B
by MiniMax · 204,800 ctx

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digita...

MiniMax: MiniMax M2.7

230B
by MiniMax · 204,800 ctx

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built...

MiniMax: MiniMax-01

457B
by MiniMax · 1,000,192 ctx

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, wi...

Minimax M1 40K

text
by MiniMax · 1,048,576 ctx

Minimax M1 80K

text
by MiniMax · 1,048,576 ctx

Ministral 3 14B Instruct 2512

14B
by Mistral AI · 262,144 ctx

Mistral (7B) Instruct v0.3

7B
by Mistral AI · 32,768 ctx

Mistral Large

text
by Mistral AI · 128,000 ctx

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excel...

Mistral Large 2407

text
by Mistral AI · 131,072 ctx

This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels ...

Mistral-Nemo-Instruct-2407

text
by Mistral AI · 131,072 ctx

12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.

Mistral-Small-3.2-24B-Instruct-2506

24B
by Mistral AI · 128,000 ctx

Mistral-Small-3.2-24B-Instruct is a drop-in upgrade over the 3.1 release, with markedly better instruction following, roughly half the in...

Mistral: Codestral 2508

text
by Mistral AI · 256,000 ctx

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks su...

Mistral: Devstral 2 2512

128B
by Mistral AI · 262,144 ctx

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer...

Mistral: Devstral Medium

24B
by Mistral AI · 131,072 ctx

Devstral Medium is a high-performance code generation and agentic reasoning model developed jointly by Mistral AI and All Hands AI. Posit...

Mistral: Devstral Small 1.1

24B
by Mistral AI · 131,072 ctx

Devstral Small 1.1 is a 24B parameter open-weight language model for software engineering agents, developed by Mistral AI in collaboratio...

Mistral: Ministral 3 14B 2512

14B
by Mistral AI · 262,144 ctx

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistra...

Mistral: Ministral 3 3B 2512

3B
by Mistral AI · 131,072 ctx

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

Mistral: Ministral 3 8B 2512

8B
by Mistral AI · 262,144 ctx

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

Mistral: Mistral 7B Instruct v0.1

7B
by Mistral AI · 4,096 ctx

A 7.3B parameter model that outperforms Llama 2 13B on all benchmarks, with optimizations for speed and context length.

Mistral: Mistral Large 3 2512

multimodal
by Mistral AI · 262,144 ctx

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active paramete...

Mistral: Mistral Medium 3

multimodal
by Mistral AI · 131,072 ctx

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly r...

Mistral: Mistral Medium 3.1

multimodal
by Mistral AI · 131,072 ctx

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to del...

Mistral: Mistral Medium 3.5

multimodal
by Mistral AI · 262,144 ctx

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and i...

Mistral: Mistral Small 3

24B
by Mistral AI · 32,768 ctx

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache...

Mistral: Mistral Small 3.1 24B

24B
by Mistral AI · 128,000 ctx

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal...

Mistral: Mistral Small 3.2 24B

24B
by Mistral AI · 128,000 ctx

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduct...

Mistral: Mistral Small 4

121B
by Mistral AI · 262,144 ctx

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into ...

Mistral: Pixtral Large 2411

multimodal
by Mistral AI · 131,072 ctx

Pixtral Large is a 124B parameter, open-weight, multimodal model built on top of [Mistral Large 2](/mistralai/mistral-large-2411). The mo...

Mistral: Saba

text
by Mistral AI · 32,768 ctx

Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextu...

Mixtral 8X22b Instruct V0.1

22B
by Mistral AI · 65,536 ctx

Mixtral 8X7b V0.1

7B
by Mistral AI · 32,768 ctx

Mixtral-8x7B Instruct v0.1

7B
by Mistral AI · 32,768 ctx

Molmo 7B D 0924

7B
by Allen Institute for AI (AI2) · 4,096 ctx
?

MoonshotAI: Kimi K2 0905

1000B
by Moonshot AI · 262,144 ctx

Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model d...

?

Morph: Morph V3 Fast

7B
by Morph · 81,920 ctx

Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the p...

?

Morph: Morph V3 Large

70B
by Morph · 262,144 ctx

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model...

?

MythoMax 13B

13B
by Gryphe · 4,096 ctx

One of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay. #merge

NVIDIA-Nemotron-3-Super-120B-A12B

120B
by Nvidia · 262,144 ctx

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent ...

NVIDIA: Llama 3.3 Nemotron Super 49B V1.5

49B
by Nvidia · 131,072 ctx

Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct wit...

NVIDIA: Nemotron 3 Nano 30B A3B

30B
by Nvidia · 262,144 ctx

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build special...

NVIDIA: Nemotron 3 Super

120B
by Nvidia · 1,000,000 ctx

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accu...

NVIDIA: Nemotron Nano 9B V2

9B
by Nvidia · 131,072 ctx

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reaso...

Nemotron-3-Nano-Omni-30B-A3B-Reasoning

30B
by Nvidia · 262,144 ctx

Nemotron 3 Nano Omni is an open multimodal model built on a hybrid Mixture-of-Experts (MoE) architecture, engineered for high efficiency ...

?

Nex AGI: DeepSeek V3.1 Nex N1

text
by Nex Agi · 131,072 ctx

DeepSeek V3.1 Nex-N1 is the flagship release of the Nex-N1 series — a post-trained model designed to highlight agent autonomy, tool use, ...

?

Nous Hermes 2 Mixtral 8X7B Dpo

7B
by Nous Research · 32,768 ctx
?

Nous: Hermes 3 405B Instruct

405B
by Nous Research · 131,072 ctx

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better rolepl...

?

Nous: Hermes 4 405B

405B
by Nous Research · 131,072 ctx

Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mo...

?

Nous: Hermes 4 70B

70B
by Nous Research · 131,072 ctx

Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the large...

?

NousResearch: Hermes 2 Pro - Llama-3 8B

8B
by Nous Research · 8,192 ctx

Hermes 2 Pro is an upgraded, retrained version of Nous Hermes 2, consisting of an updated and cleaned version of the OpenHermes 2.5 Datas...

Nvidia Nemotron Nano 9B V2

9B
by Nvidia · 131,072 ctx
?

Perceptron: Perceptron Mk1

multimodal
by Perceptron · 32,768 ctx

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and ...

?

Prime Intellect: INTELLECT-3

text
by Prime Intellect · 131,072 ctx

INTELLECT-3 is a 106B-parameter Mixture-of-Experts model (12B active) post-trained from GLM-4.5-Air-Base using supervised fine-tuning (SF...

Qwen 2 (1.5B)

5B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2 (72B)

72B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2 (7B)

7B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2 Instruct (1.5B)

5B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2.5 14B Instruct

14B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2.5 Coder 32B Instruct

32B
by Alibaba (Qwen Team) · 16,384 ctx

Qwen QwQ-32B

32B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2 72B Instruct

72B
by Togethercomputer · 32,768 ctx

Qwen2-VL (72B) Instruct

72B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 1.5B

5B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 1.5B Instruct

5B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 14B

14B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 32B

32B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 32B Instruct

32B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 3B Instruct

3B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 72B

72B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 72B Instruct

72B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 72B Instruct Turbo

72B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 7B

7B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 7B Instruct

7B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 7B Instruct Turbo

7B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 0.6B

6B
by Alibaba (Qwen Team) · 40,960 ctx

Qwen3 0.6B Base

6B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 1.7B

7B
by Alibaba (Qwen Team) · 40,960 ctx

Qwen3 1.7B Base

7B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 14B Base

14B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 30B A3b Base

30B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 4B Base

4B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 4B Instruct 2507

4B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 8B Base

8B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3-235B-A22B-Instruct-2507

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-235B-A22B-Instruct-2507 is the updated version of the Qwen3-235B-A22B non-thinking mode, featuring Significant improvements in gene...

Qwen3.5-0.8B

8B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5-0.8B is Alibaba's smallest model in the Qwen3.5 series, featuring a hybrid Gated Delta Networks and sparse Mixture-of-Experts arc...

Qwen3.5-2B

2B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5-2B is a compact yet capable model from Alibaba's Qwen3.5 series. It features a 262K token context window, support for 201 languag...

Qwen3.5-4B

4B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5-4B is a mid-size model from Alibaba's Qwen3.5 series that delivers a strong balance of performance and efficiency. It features a ...

Qwen: Qwen Plus 0728

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, an...

Qwen: Qwen-Plus

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

Qwen: Qwen2.5 7B Instruct

7B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more...

Qwen: Qwen2.5 VL 72B Instruct

72B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing ...

Qwen: Qwen3 14B

14B
by Alibaba (Qwen Team) · 131,702 ctx

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialo...

Qwen: Qwen3 235B A22B

235B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supp...

Qwen: Qwen3 235B A22B Instruct 2507

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture...

Qwen: Qwen3 235B A22B Thinking 2507

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning ...

Qwen: Qwen3 30B A3B

30B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to e...

Qwen: Qwen3 30B A3B Instruct 2507

30B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. ...

Qwen: Qwen3 30B A3B Thinking 2507

30B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-st...

Qwen: Qwen3 32B

32B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dial...

Qwen: Qwen3 8B

8B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dia...

Qwen: Qwen3 Coder 30B A3B Instruct

30B
by Alibaba (Qwen Team) · 160,000 ctx

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed f...

Qwen: Qwen3 Coder 480B A35B

480B
by Alibaba (Qwen Team) · 1,048,576 ctx

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agenti...

Qwen: Qwen3 Coder Flash

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model...

Qwen: Qwen3 Coder Next

480B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse Mo...

Qwen: Qwen3 Coder Plus

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializ...

Qwen: Qwen3 Max

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual ...

Qwen: Qwen3 Max Thinking

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi...

Qwen: Qwen3 Next 80B A3B Instruct

80B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thi...

Qwen: Qwen3 Next 80B A3B Thinking

80B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. ...

Qwen: Qwen3 VL 235B A22B Instruct

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across image...

Qwen: Qwen3 VL 235B A22B Thinking

235B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. ...

Qwen: Qwen3 VL 30B A3B Instruct

30B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its ...

Qwen: Qwen3 VL 30B A3B Thinking

30B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its ...

Qwen: Qwen3 VL 32B Instruct

32B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across te...

Qwen: Qwen3 VL 8B Instruct

8B
by Alibaba (Qwen Team) · 256,000 ctx

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning ...

Qwen: Qwen3 VL 8B Thinking

8B
by Alibaba (Qwen Team) · 256,000 ctx

Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual rea...

Qwen: Qwen3.5 397B A17B

397B
by Alibaba (Qwen Team) · 262,144 ctx

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism ...

Qwen: Qwen3.5 Plus 2026-02-15

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with...

Qwen: Qwen3.5 Plus 2026-04-20

235B
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces t...

Qwen: Qwen3.5-122B-A10B

122B
by Alibaba (Qwen Team) · 262,144 ctx

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a ...

Qwen: Qwen3.5-27B

27B
by Alibaba (Qwen Team) · 262,144 ctx

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balanc...

Qwen: Qwen3.5-35B-A3B

35B
by Alibaba (Qwen Team) · 262,144 ctx

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechani...

Qwen: Qwen3.5-9B

9B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understandi...

Qwen: Qwen3.5-Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sp...

Qwen: Qwen3.6 27B

27B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid mult...

Qwen: Qwen3.6 35B A3B

35B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters pe...

Qwen: Qwen3.6 Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M toke...

Qwen: Qwen3.6 Plus

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling s...

Qwen: Qwen3.7 Max

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric worklo...

?

ReMM SLERP 13B

13B
by Undi95 · 6,144 ctx

A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge

Reka Edge

7B
by Rekaai · 16,384 ctx

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. ...

Reka Flash 3

21B
by Rekaai · 65,536 ctx

Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at gen...

?

Relace: Relace Apply 3

text
by Relace · 256,000 ctx

Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates fr...

?

Relace: Relace Search

text
by Relace · 256,000 ctx

The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user re...

?

Sao10K: Llama 3 8B Lunaris

8B
by Sao10k · 8,192 ctx

Lunaris 8B is a versatile generalist and roleplaying model based on Llama 3. It's a strategic merge of multiple models, designed to balan...

?

Sao10K: Llama 3.1 70B Hanami x1

70B
by Sao10k · 16,000 ctx

This is [Sao10K](/sao10k)'s experiment over [Euryale v2.2](/sao10k/l3.1-euryale-70b).

?

Sao10K: Llama 3.1 Euryale 70B v2.2

70B
by Sao10k · 131,072 ctx

Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3...

?

Sao10K: Llama 3.3 Euryale 70B

70B
by Sao10k · 131,072 ctx

Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B ...

?

Sao10k: Llama 3 Euryale 70B v2.1

70B
by Sao10k · 8,192 ctx

Euryale 70B v2.1 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). - Better prompt adherence. - Better ana...

Sarvam M

24B
by Sarvamai · 32,768 ctx
?

Seed-1.8

200B
by Bytedance · 256,000 ctx

Optimized specifically for multimodal agent scenarios. It features enhanced agent capabilities, upgraded multimodal comprehension, and mo...

?

Seed-2.0-code

multimodal
by Bytedance · 256,000 ctx

A coding model optimized for real-world development environments, with reliable tool use in common IDEs such as Claude Code. It delivers ...

?

Seed-2.0-pro

multimodal
by Bytedance · 256,000 ctx

Built for the Agent era, it delivers stable performance in complex reasoning and long-horizon tasks, including multi-step planning, visua...

StepFun: Step 3.5 Flash

199B
by Stepfun · 262,144 ctx

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it select...

?

Switchpoint Router

text
by Switchpoint · 131,072 ctx

Switchpoint AI's router instantly analyzes your request and directs it to the optimal AI from an ever-evolving library. As the world of L...

Tencent: Hunyuan A13B Instruct

80B
by Tencent · 131,072 ctx

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B ...

?

TheDrummer: Cydonia 24B V4.1

24B
by Thedrummer · 131,072 ctx

Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.

?

TheDrummer: Rocinante 12B

12B
by Thedrummer · 32,768 ctx

Rocinante 12B is designed for engaging storytelling and rich prose. Early testers have reported: - Expanded vocabulary with unique and ex...

?

TheDrummer: Skyfall 36B V2

36B
by Thedrummer · 32,768 ctx

Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-pla...

?

TheDrummer: UnslopNemo 12B

12B
by Thedrummer · 32,768 ctx

UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.

Tongyi DeepResearch 30B A3B

30B
by Alibaba (Qwen Team) · 131,072 ctx

Tongyi DeepResearch is an agentic large language model developed by Tongyi Lab, with 30 billion total parameters activating only 3 billio...

Upstage: Solar Pro 3

text
by Upstage · 128,000 ctx

Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forwa...

WizardLM-2 8x22B

22B
by Microsoft · 65,536 ctx

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprie...

Writer: Palmyra X5

70B
by Writer · 1,040,000 ctx

Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-...

Xiaomi: MiMo-V2-Flash

7B
by Xiaomi · 262,144 ctx

MiMo-V2-Flash is an open-source foundation language model developed by Xiaomi. It is a Mixture-of-Experts model with 309B total parameter...

Xiaomi: MiMo-V2-Omni

multimodal
by Xiaomi · 262,144 ctx

MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It comb...

Xiaomi: MiMo-V2-Pro

text
by Xiaomi · 1,048,576 ctx

MiMo-V2-Pro is Xiaomi's flagship foundation model, featuring over 1T total parameters and a 1M context length, deeply optimized for agent...

Xiaomi: MiMo-V2.5

315B
by Xiaomi · 1,048,576 ctx

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surp...

Xiaomi: MiMo-V2.5-Pro

65B
by Xiaomi · 1,048,576 ctx

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, an...

Z.ai: GLM 4 32B

32B
by Zhipu AI · 128,000 ctx

GLM 4 32B is a cost-effective foundation language model. It can efficiently perform complex tasks and has significantly enhanced capabili...

Z.ai: GLM 4.5V

9B
by Zhipu AI · 65,536 ctx

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 1...

Z.ai: GLM 4.6V

108B
by Zhipu AI · 131,072 ctx

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents,...

Z.ai: GLM 4.7 Flash

31B
by Zhipu AI · 202,752 ctx

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agenti...

Z.ai: GLM 5V Turbo

multimodal
by Zhipu AI · 202,752 ctx

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively ...

gemini-3.1-pro

multimodal
by Google DeepMind · 1,000,000 ctx

Bring any idea to life with state-of-the-art reasoning to help you learn, build, and plan anything. Best for complex tasks and bringing c...

?

inclusionAI: Ling-2.6-1T

text
by Inclusionai · 262,144 ctx

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents...

?

inclusionAI: Ling-2.6-flash

text
by Inclusionai · 262,144 ctx

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-w...

?

inclusionAI: Ring-2.6-1T

text
by Inclusionai · 262,144 ctx

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both str...

meta-llama/Llama-2-7b-chat-hf

7B
by Meta AI · 4,096 ctx

meta-llama/Llama-3.3-70B-Instruct

70B
by Meta AI · 131,072 ctx

nim/meta/llama-3.1-70b-instruct

70B
by Meta AI · 16,384 ctx

nim/meta/llama-3.1-8b-instruct

8B
by Meta AI · 16,384 ctx

nim/meta/llama-3.2-11b-vision-instruct

11B
by Nvidia · 16,384 ctx

nim/meta/llama-3.2-90b-vision-instruct

90B
by Meta AI · 16,384 ctx

nim/meta/llama-3.3-70b-instruct

70B
by Meta AI · 16,384 ctx

nim/mistralai/mixtral-8x22b-instruct-v01

22B
by Mistral AI · 16,384 ctx

nim/mistralai/mixtral-8x7b-instruct-v01

7B
by Mistral AI · 16,384 ctx

nim/nv-mistralai/mistral-nemo-12b-instruct

12B
by Nvidia · 16,384 ctx

nim/nvidia/llama-3.1-nemotron-70b-instruct

70B
by Nvidia · 16,384 ctx

nim/nvidia/llama-3.3-nemotron-super-49b-v1

49B
by Nvidia · 16,384 ctx
?

Kimi K2.6

1000B
by Moonshot AI · Kimi · 256,000 ctx

Long-horizon coding + autonomous-execution upgrade over K2.5.

?

Kimi K2.5

1000B
by Moonshot AI · Kimi · 256,000 ctx

Multimodal agentic variant — adds a vision encoder to the K2 backbone.

?

Kimi K2 Thinking

1000B
by Moonshot AI · Kimi · 256,000 ctx

Moonshot's open-weight reasoning variant — extended chain-of-thought training on top of Kimi K2.

GPT-OSS 120B

120B
by OpenAI · GPT-OSS · 128,000 ctx

OpenAI's first open-weight LLM in years — Apache-licensed MoE.

GPT-OSS 20B

20B
by OpenAI · GPT-OSS · 128,000 ctx

20B GPT-OSS — single-GPU local target.

GLM-4.5

355B
by Zhipu AI · GLM · 128,000 ctx

Zhipu's frontier open-weight MoE — 355B total, 32B active. Strong agentic + reasoning marks for an open model.

GLM-4.5-Air

106B
by Zhipu AI · GLM · 128,000 ctx

Smaller, cheaper sibling of GLM-4.5. 106B total, 12B active.

?

Kimi K2

1000B
by Moonshot AI · Kimi · 256,000 ctx

Moonshot's frontier open-weight MoE — 1T total, 32B active.

Qwen 3 14B

15B
by Alibaba (Qwen Team) · Qwen 3 · 128,000 ctx

Qwen 3 235B

235B
by Alibaba (Qwen Team) · Qwen 3 · 128,000 ctx

Alibaba's frontier MoE — 235B total / 22B active.

Qwen 3 32B

33B
by Alibaba (Qwen Team) · Qwen 3 · 128,000 ctx

Dense 32B Qwen 3.

Qwen 3 4B

4B
by Alibaba (Qwen Team) · Qwen 3 · 32,768 ctx

Qwen 3 8B

8B
by Alibaba (Qwen Team) · Qwen 3 · 128,000 ctx

Gemma 3 12B

12B
by Google DeepMind · Gemma · 128,000 ctx

12B Gemma 3 — multimodal, single-GPU target.

Gemma 3 1B

1B
by Google DeepMind · Gemma · 32,768 ctx

1B Gemma 3 — edge / mobile.

Gemma 3 27B

27B
by Google DeepMind · Gemma · 128,000 ctx

Google's open-weight multimodal LLM — efficient and license-permissive.

Gemma 3 4B

4B
by Google DeepMind · Gemma · 128,000 ctx

4B Gemma 3 — laptop multimodal.

DeepSeek R1

671B
by DeepSeek · DeepSeek · 128,000 ctx

DeepSeek's reasoning model — RL-trained, frontier-class, MIT-licensed.

DeepSeek R1 Distill Llama 70B

70B
by DeepSeek · DeepSeek · 128,000 ctx

70B Llama distilled from DeepSeek R1's reasoning traces.

DeepSeek R1 Distill Qwen 1.5B

2B
by DeepSeek · DeepSeek · 128,000 ctx

Tiny distilled R1 — phone / browser deployable.

DeepSeek R1 Distill Qwen 14B

15B
by DeepSeek · DeepSeek · 128,000 ctx

14B distilled R1 — laptop-friendly reasoning.

DeepSeek R1 Distill Qwen 32B

33B
by DeepSeek · DeepSeek · 128,000 ctx

32B Qwen base distilled from DeepSeek R1.

DeepSeek R1 Distill Qwen 7B

8B
by DeepSeek · DeepSeek · 128,000 ctx

7B distilled R1 — runs on any modern GPU.

MiniMax-Text-01

456B
by MiniMax · MiniMax-Text · 4,000,000 ctx

MiniMax's first open MoE — 456B total, 45.9B active. 1M+ context via Lightning Attention.

OLMo 3 7B

7B
by Allen Institute for AI (AI2) · OLMo · 4,096 ctx

Allen AI's latest fully-open OLMo — model + training data + checkpoints.

DeepSeek V3

671B
by DeepSeek · DeepSeek · 128,000 ctx

DeepSeek's flagship MoE — 671B total, 37B active, frontier-class.

IBM Granite 3.1 8B

8B
by IBM Research · Granite · 128,000 ctx

IBM's enterprise-focused open-weight LLM.

Phi-4

15B
by Microsoft · Phi · 16,384 ctx

Microsoft's 14B small-LM workhorse — punches above its weight.

Llama 3.3 70B

70B
by Meta AI · Llama · 128,000 ctx

Meta's best-in-class open-weight LLM — 70B class.

Qwen 2.5 Coder 32B

33B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

Alibaba's open-weight coding model — best in class for 32B.

Hunyuan-Large

389B
by Tencent · Hunyuan · 256,000 ctx

Tencent's open-weight MoE — 389B total, 52B active. Largest open MoE at launch.

Stable Diffusion 3.5 Medium

3B
by Stability AI · Stable Diffusion

2.5B SD 3.5 — fits on 12 GB consumer GPUs.

Stable Diffusion 3.5 Large

8B
by Stability AI · Stable Diffusion

Stability AI's latest open-weight image-gen — 8.1B params, MMDiT architecture.

Yi-Lightning

text
by 01.AI · Yi · 16,000 ctx

01.AI's fastest production model. Tops the LMSYS Arena Chinese leaderboard.

Llama 3.2 11B Vision

11B
by Meta AI · Llama · 128,000 ctx

Meta's open-weight multimodal LLM — vision + text in 11B.

Llama 3.2 1B

1B
by Meta AI · Llama · 128,000 ctx

Meta's smallest Llama — mobile + on-device target.

Llama 3.2 3B

3B
by Meta AI · Llama · 128,000 ctx

3B Llama — laptop-class chat + RAG.

Llama 3.2 90B Vision

90B
by Meta AI · Llama · 128,000 ctx

Meta's largest vision-capable Llama.

Qwen 2.5 14B

15B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

14B Qwen 2.5 — sweet spot for single-GPU local hosting.

Qwen 2.5 32B

33B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

32B Qwen 2.5 — laptop-class workhorse.

Qwen 2.5 3B

3B
by Alibaba (Qwen Team) · Qwen · 32,768 ctx

3B Qwen 2.5 — laptop / edge target.

Qwen 2.5 72B

73B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

Alibaba's flagship open-weight LLM — 72B dense.

Qwen 2.5 7B

8B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

7B Qwen 2.5 — most popular Qwen variant on Ollama.

Command R+

104B
by Cohere · Command · 128,000 ctx

Cohere's open-weight RAG-optimized LLM — multilingual + tool use.

Phi-3.5 Mini

4B
by Microsoft · Phi · 128,000 ctx

3.8B Phi — laptop / edge target.

?

Hermes 3 70B

70B
by Nous Research · Hermes · 128,000 ctx

Nous Research's flagship Llama fine-tune — agent-friendly.

?

Hermes 3 8B

8B
by Nous Research · Hermes · 128,000 ctx

FLUX.1 Dev

12B
by Black Forest Labs · FLUX

Open-weight FLUX.1 — non-commercial license.

FLUX.1 Pro

12B
by Black Forest Labs · FLUX

Black Forest Labs' flagship image-gen model — closed/API.

FLUX.1 Schnell

12B
by Black Forest Labs · FLUX

Distilled fast FLUX.1 — Apache-2.0, commercial-friendly.

Gemma 2 2B

3B
by Google DeepMind · Gemma 2 · 8,192 ctx

Tiny 2B Gemma 2 — laptop / mobile.

Mistral Large 2

123B
by Mistral AI · Mistral Large · 128,000 ctx

Mistral's flagship open-weight model — 123B dense.

Llama 3.1 405B

405B
by Meta AI · Llama · 128,000 ctx

Meta's largest open-weight LLM — dense 405B, frontier-class at launch.

Llama 3.1 70B

70B
by Meta AI · Llama · 128,000 ctx

Llama 3.1 70B — production workhorse, superseded by 3.3 but still widely deployed.

Llama 3.1 8B

8B
by Meta AI · Llama · 128,000 ctx

Meta's most popular open-weight small LLM — fits anywhere.

Mistral Nemo 12B

12B
by Mistral AI · Mistral · 128,000 ctx

Mistral × Nvidia collab — 12B Apache-licensed, multilingual.

InternLM 2.5 20B

20B
by Shanghai AI Lab · InternLM · 1,000,000 ctx

Shanghai AI Lab's dense 20B open-weight. Strong long-context + tool use for its size.

Gemma 2 27B

27B
by Google DeepMind · Gemma 2 · 8,192 ctx

Google's pre-Gemma-3 open-weight workhorse.

Gemma 2 9B

9B
by Google DeepMind · Gemma 2 · 8,192 ctx

9B Gemma 2 — single-GPU local target.

DeepSeek Coder V2 236B

236B
by DeepSeek · DeepSeek Coder · 128,000 ctx

DeepSeek's MoE coding model — 236B total, 21B active.

DeepSeek Coder V2 Lite

16B
by DeepSeek · DeepSeek Coder · 128,000 ctx

16B MoE / 2.4B active — laptop-class coder.

Mistral 7B v0.3

7B
by Mistral AI · Mistral 7B · 32,768 ctx

The current Mistral 7B — adds function calling + extended vocab.

IBM Granite Code 8B

8B
by IBM Research · Granite · 4,096 ctx

Phi-3 Medium

14B
by Microsoft · Phi · 128,000 ctx

Phi-3 Mini

4B
by Microsoft · Phi · 128,000 ctx

Mixtral 8x22B

141B
by Mistral AI · Mixtral · 65,536 ctx

Mistral's open MoE — 141B total, 39B active.

Mistral 7B v0.2

7B
by Mistral AI · Mistral 7B · 32,768 ctx

Mistral 7B v0.2 — earlier 32K context revision.

mxbai-embed-large

0B
by Mixedbread AI · Mixedbread Embed · 512 ctx

335M embedding model — top MTEB scores for its size.

Moondream 1.8B

2B
by Moondream · Moondream · 2,048 ctx

Tiny multimodal — laptop-class image understanding.

Nomic Embed Text

0B
by Nomic AI · Nomic Embed · 8,192 ctx

Open embedding model — 69M Ollama pulls, the local default.

OLMo 7B

7B
by Allen Institute for AI (AI2) · OLMo · 2,048 ctx

LLaVA 13B

13B
by LLaVA Project · LLaVA · 4,096 ctx

LLaVA 34B

34B
by LLaVA Project · LLaVA · 4,096 ctx

Largest open-weight LLaVA — vision encoder + Yi-34B backbone.

LLaVA 7B

7B
by LLaVA Project · LLaVA · 4,096 ctx
?

BGE-M3

1B
by BAAI (Beijing Academy of AI) · BGE · 8,192 ctx

Multilingual + multifunctional embedding (100+ languages).

Code Llama 70B

70B
by Meta AI · Code Llama · 16,384 ctx

Meta's largest code-specialised Llama.

TinyLlama 1.1B

1B
by TinyLlama Project · TinyLlama · 2,048 ctx

1.1B Llama-arch model — 3T training tokens.

Whisper Large v3

2B
by OpenAI · Whisper · 30 ctx

OpenAI's open-weight speech-to-text — the standard transcription model.

DeepSeek Coder 33B

33B
by DeepSeek · DeepSeek Coder · 16,384 ctx

DeepSeek Coder 6.7B

7B
by DeepSeek · DeepSeek Coder · 16,384 ctx

Yi-34B

34B
by 01.AI · Yi · 32,000 ctx

Earlier open-weight Yi release — bilingual EN/ZH, 32K-token context.

Mistral 7B v0.1

7B
by Mistral AI · Mistral 7B · 8,192 ctx

Original Mistral 7B — historical reference.

Baichuan2-13B

13B
by Baichuan Inc. · Baichuan · 4,096 ctx

Bilingual EN/ZH open-weight. Strong for its size on Chinese-language benchmarks.

Code Llama 13B

13B
by Meta AI · Code Llama · 16,384 ctx

Code Llama 34B

34B
by Meta AI · Code Llama · 16,384 ctx

Code Llama 7B

7B
by Meta AI · Code Llama · 16,384 ctx

Stable Diffusion XL

4B
by Stability AI · Stable Diffusion

Workhorse open-weight image-gen — 3.5B params, runs anywhere.

Whisper Base

0B
by OpenAI · Whisper · 30 ctx

74M Whisper — browser / Raspberry Pi-deployable.

Whisper Medium

1B
by OpenAI · Whisper · 30 ctx

769M Whisper variant — half the size of Large, 80% of the accuracy.

Whisper Small

0B
by OpenAI · Whisper · 30 ctx

244M Whisper — fits on edge GPUs and CPU.

Whisper Tiny

0B
by OpenAI · Whisper · 30 ctx

39M Whisper — runs in-browser via WebGPU.

Stable Diffusion 1.5

1B
by Stability AI · Stable Diffusion

The original viral image-gen model — still searched heavily.

Meta: Muse Spark 1.3 Contributor

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and...

GLM 5.3 FP8

text
by Zhipu AI · 1,048,576 ctx

GLM 5.3 FP8 Lora

text
by Zhipu AI · 1,048,576 ctx
?

Inception: Mercury 2.5 Preview

text
by Inception · 260,000 ctx

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, ...

Tencent: Hy4 preview

text
by Tencent · 1,048,576 ctx

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding ...

DeepSeek: DeepSeek V4 Flash Vision Exp

multimodal
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepsee...

Meta: Muse Spark 1.2 Contributor

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. ...

GLM 5.2 FP8 Lora

text
by Zhipu AI · 1,048,576 ctx

GLM 5.2 FP8

text
by Zhipu AI · 1,048,576 ctx

gemma-4-31B-it-Ultra

31B
by Google DeepMind · 131,072 ctx

Ultra speed version of gemma-4-31B-it

gpt-oss-120b-Ultra

120B
by OpenAI · 131,072 ctx

Ultra fast version of gpt-oss-120b

?

Inkling FP4

952B
by Thinking Machines · 524,288 ctx

Qwen3.5 0.8B Lora

8B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 122B A10B Lora

122B
by Alibaba (Qwen Team) · 8,192 ctx

Qwen3.5 27B Lora

27B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 2B Lora

2B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 35B A3B Base Lora

35B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 397B A17B Lora

397B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 4B Lora

4B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 9B Lora

9B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6 27B Lora

27B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6 35B A3B Lora

35B
by Alibaba (Qwen Team) · 262,144 ctx
?

MiniMax-M2.7-Turbo

text
by Minimaxai · 196,608 ctx

Speed-optimized MiniMax-M2.7

NVIDIA Nemotron 3 Ultra NVFP4

text
by Nvidia · 262,144 ctx

Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, re...

NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16

550B
by Nvidia · 262,144 ctx

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It deli...

Qwen3.5 35B A3b LoRa

35B
by Alibaba (Qwen Team) · 262,144 ctx

GLM 5 Fp4

text
by Zhipu AI · 202,752 ctx

Glm 4.7 Fp8

text
by Zhipu AI · 202,752 ctx

Llama 4 Maverick 17B 128E Instruct Nvfp4

17B
by Meta AI · 1,048,576 ctx

GLM 4.7 FP4

text
by Togethercomputer · 202,752 ctx

Kimi K2.5 FP4

text
by Togethercomputer · 262,144 ctx

MiniMax M2.5 FP4

text
by MiniMax · 8,192 ctx

Cogito V1 Preview Llama 70B

70B
by Deepcogito · 131,072 ctx

Cogito V1 Preview Llama 70B Turbo

70B
by Deepcogito · 131,072 ctx

Cogito V1 Preview Llama 8B

8B
by Deepcogito · 131,072 ctx

Cogito V1 Preview Qwen 14B

14B
by Deepcogito · 131,072 ctx

Cogito V1 Preview Qwen 32B

32B
by Deepcogito · 131,072 ctx

DeepSeek: DeepSeek V3.2 Exp

671B
by DeepSeek · 163,840 ctx

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectu...

Deepcoder 14B Preview

14B
by Togethercomputer · 131,072 ctx

Gemma 3 270M It Lora

0B
by Google DeepMind · 32,768 ctx

Gemma 3 27B It Lora

27B
by Google DeepMind

Gemma 4 31B It Lora

31B
by Google DeepMind · 262,144 ctx

Glm 4.5 Air Fp8

text
by Zhipu AI · 131,072 ctx
?

L3-8B-Lunaris-v1-Turbo

8B
by Sao10k · 8,192 ctx

Llama 3.3 70B Instruct FP8 Lora

70B
by Meta AI · 131,072 ctx

Llama 4 Maverick Instruct (17Bx128E) FP8

17B
by Meta AI · 1,048,576 ctx

Llama 4 Scout 17B 16E Instruct Fp8 Lora

17B
by Meta AI · 10,485,760 ctx

Meta Llama 3.1 8B Instruct Awq Int4

8B
by Meta AI · 131,072 ctx

Mixtral 8x7B Instruct V0.1 FP8 Lora

7B
by Mistral AI · 32,768 ctx
?

MoonshotAI Kimi Latest

1000B
by Moonshot AI · 262,144 ctx

This model always redirects to the latest model in the MoonshotAI Kimi family.

Nemotron 3 Nano Omni 30B A3b Reasoning Fp8

30B
by Nvidia · 131,072 ctx

Nvidia Nemotron 3 Nano 30B A3b Bf16

30B
by Nvidia · 262,144 ctx

Nvidia Nemotron 3 Super 120B A12b Bf16

120B
by Nvidia · 262,144 ctx

Nvidia Nemotron 3 Super 120B A12b Fp8

120B
by Nvidia · 262,144 ctx

Qwen3 235B A22B Instruct 2507 FP8 Throughput

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 30B A3B Instruct 2507 Lora

30B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 8B Lora

8B
by Alibaba (Qwen Team) · 40,960 ctx

Qwen3 Coder 480B A35B Instruct Fp8

480B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 Coder Next Fp8

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 Next 80B A3b Instruct Fp8

80B
by Alibaba (Qwen Team)

Qwen3-Coder-480B-A35B-Instruct-Turbo

480B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Coder-480B-A35B-Instruct is the Qwen3's most agentic code model, featuring Significant Performance on Agentic Coding, Agentic Brows...

Qwen3-VL-235B-A22B-Instruct-FP8

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 122B A10b Fp8

122B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 9B Fp8

9B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6 35B A3b Fp8

35B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen: Qwen3.6 Max Preview

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximate...

Tencent: Hy3 preview

299B
by Tencent · 262,144 ctx

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports con...

gemma-4-31B-it-turbo

31B
by Google DeepMind · 262,144 ctx

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating te...

gpt-oss-120b-Turbo

120B
by OpenAI · 131,072 ctx

Closed / API-only models.

Direct API, aggregator (OpenRouter, Bedrock), or chat UI.

Google: Gemini 3.8 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic task...

Anthropic: Claude Fable 5.1

multimodal
by Anthropic · 1,000,000 ctx

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, a...

?

Ox Alpha

multimodal
by Stealth · 1,048,576 ctx

Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon softwa...

Google: Gemini 3.7 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed f...

SpaceXAI: Grok 4.6

multimodal
by xAI · 500,000 ctx

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

?

Sakana: Sakana Namazu

multimodal
by Sakana · 262,144 ctx

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language...

?

Thinking Machines: Inkling Small

multimodal
by Thinkingmachines · 524,288 ctx

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B to...

Claude Opus 5

multimodal
by Anthropic · 1,000,000 ctx

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at ...

Google: Gemini 3.5 Flash-Lite

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.5 Flash-Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute ...

Google: Gemini 3.6 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to pro...

?

Meituan: LongCat 2.0

text
by Meituan · 1,048,756 ctx

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for co...

?

Poolside: Laguna S 2.1

text
by Poolside · 1,048,576 ctx

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with ...

?

Auto Router (Beta)

multimodal
by Openrouter · 2,000,000 ctx

Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#...

OpenAI: GPT-5.6 Luna

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as ch...

OpenAI: GPT-5.6 Luna Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` ...

OpenAI: GPT-5.6 Sol

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is p...

OpenAI: GPT-5.6 Sol Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set...

OpenAI: GPT-5.6 Terra

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. ...

OpenAI: GPT-5.6 Terra Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mod...

?

Venice: Uncensored

24B
by Cognitivecomputations · 128,000 ctx

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in ...

xAI: Grok 4.5

multimodal
by xAI · 500,000 ctx

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

?

Poolside: Laguna XS 2.1

text
by Poolside · 262,144 ctx

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from thei...

Anthropic: Claude Sonnet 5

multimodal
by Anthropic · 1,000,000 ctx

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It suppo...

Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

multimodal
by Google DeepMind · 65,536 ctx

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity dev...

?

Sakana: Fugu Ultra

multimodal
by Sakana · 1,000,000 ctx

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-age...

?

Poolside: Laguna M.1

text
by Poolside · 262,144 ctx

Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. De...

?

Poolside: Laguna XS.2

text
by Poolside · 262,144 ctx

Laguna XS.2 is the second-generation model in the XS size class from [Poolside](https://poolside.ai/), their efficient coding agent serie...

Google: Nano Banana 2 (Gemini 3.1 Flash Image)

multimodal
by Google DeepMind · 131,072 ctx

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-le...

Google: Nano Banana Pro (Gemini 3 Pro Image)

multimodal
by Google DeepMind · 65,536 ctx

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana ...

Anthropic: Claude Fable 5

multimodal
by Anthropic · 1,000,000 ctx

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file ...

?

OpenRouter: Fusion

text
by Openrouter · 128,000 ctx

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel wit...

Anthropic: Claude Opus 4.8

multimodal
by Anthropic · 1,000,000 ctx

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with t...

AI21: Jamba Large 1.7

text
by Ai21 · 256,000 ctx

Jamba Large 1.7 is the latest model in the Jamba open family, offering improvements in grounding, instruction-following, and overall effi...

Amazon: Nova 2 Lite

multimodal
by Amazon · 1,000,000 ctx

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. ...

Amazon: Nova Lite 1.0

multimodal
by Amazon · 300,000 ctx

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to ...

Amazon: Nova Micro 1.0

text
by Amazon · 128,000 ctx

Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low c...

Amazon: Nova Premier 1.0

multimodal
by Amazon · 1,000,000 ctx

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for dis...

Amazon: Nova Pro 1.0

multimodal
by Amazon · 300,000 ctx

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide ...

Anthropic: Claude 3 Haiku

multimodal
by Anthropic · 200,000 ctx

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. S...

Anthropic: Claude Opus 4

multimodal
by Anthropic · 200,000 ctx

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-runnin...

Anthropic: Claude Opus 4.1

multimodal
by Anthropic · 200,000 ctx

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic task...

Anthropic: Claude Opus 4.5

multimodal
by Anthropic · 200,000 ctx

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon c...

Anthropic: Claude Opus 4.6

multimodal
by Anthropic · 1,000,000 ctx

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire...

Anthropic: Claude Sonnet 4

multimodal
by Anthropic · 1,000,000 ctx

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with...

Anthropic: Claude Sonnet 4.5

multimodal
by Anthropic · 1,000,000 ctx

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers st...

?

Auto Router

multimodal
by Openrouter · 2,000,000 ctx

Your prompt will be processed by a meta-model and routed to one of dozens of models (see below), optimizing for the best possible output....

?

Body Builder (beta)

text
by Openrouter · 128,000 ctx

Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI mod...

Cohere: Command A

text
by Cohere · 256,000 ctx

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, mult...

Cohere: Command R (08-2024)

text
by Cohere · 128,000 ctx

command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmente...

Cohere: Command R+ (08-2024)

text
by Cohere · 128,000 ctx

command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower l...

Cohere: Command R7B (12-2024)

text
by Cohere · 128,000 ctx

Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, an...

?

Free Models Router

multimodal
by Openrouter · 200,000 ctx

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenR...

Google: Gemini 2.0 Flash

multimodal
by Google DeepMind · 1,000,000 ctx

Gemini Flash 2.0 offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while...

Google: Gemini 2.0 Flash Lite

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), ...

Google: Gemini 2.5 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and sci...

Google: Gemini 2.5 Flash Lite

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It ...

Google: Gemini 3.1 Flash Lite

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text,...

Google: Gemini 3.5 Flash

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed....

Google: Gemma 2 27B

27B
by Google DeepMind · 8,192 ctx

Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). ...

Google: Gemma 3 12B

12B
by Google DeepMind · 131,072 ctx

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, unders...

Google: Gemma 3n 4B

4B
by Google DeepMind · 32,768 ctx

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It support...

Google: Gemma 4 26B A4B

26B
by Google DeepMind · 262,144 ctx

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B...

Google: Gemma 4 31B

31B
by Google DeepMind · 262,144 ctx

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K ...

Google: Nano Banana (Gemini 2.5 Flash Image)

multimodal
by Google DeepMind · 32,768 ctx

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual...

?

Inflection: Inflection 3 Pi

text
by Inflection · 8,000 ctx

Inflection 3 Pi powers Inflection's [Pi](https://pi.ai) chatbot, including backstory, emotional intelligence, productivity, and safety. I...

?

Inflection: Inflection 3 Productivity

text
by Inflection · 8,000 ctx

Inflection 3 Productivity is optimized for following instructions. It is better for tasks requiring JSON output or precise adherence to p...

?

AionLabs: Aion 3.5

text
by Aion Labs · 262,144 ctx

Aion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative g...

?

AionLabs: Aion 3.5 Mini

text
by Aion Labs · 262,144 ctx

Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, l...

Anthropic: Claude Haiku 5.5

multimodal
by Anthropic · 1,000,000 ctx

Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use....

Anthropic: Claude Opus 5.5

multimodal
by Anthropic · 1,000,000 ctx

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. I...

Anthropic: Claude Sonnet 5.5

multimodal
by Anthropic · 1,000,000 ctx

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is ...

Cohere: Command A+

multimodal
by Cohere · 192,000 ctx

Command A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supp...

?

DeepSeek: DeepSeek Flash Latest

multimodal
by ~deepseek · 1,048,576 ctx

This model always redirects to the latest model in the DeepSeek Flash family.

?

DeepSeek: DeepSeek Pro Latest

text
by ~deepseek · 1,048,576 ctx

This model always redirects to the latest model in the DeepSeek Pro family.

Google: Nano Banana 2.1

multimodal
by Google DeepMind · 65,536 ctx

Nano Banana 2.1 (Gemini Nano Banana 2.1) is Google's image generation and editing model on the Flash tier, succeeding Nano Banana 2 and N...

?

inclusionAI: Ling 3.1 Flash

text
by Inclusionai · 262,144 ctx

Ling 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.

?

Inference.net: Schematron V2 Small

text
by Inference Net · 128,000 ctx

Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex sch...

?

Inference.net: Schematron V2 Turbo

text
by Inference Net · 128,000 ctx

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extract...

Mistral: Mistral Large 4

multimodal
by Mistral AI · 524,288 ctx

Mistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads....

?

Nex AGI: Nex-N2.5-Mini

multimodal
by Nex Agi · 262,144 ctx

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual fee...

?

Nex AGI: Nex-N2.5-Pro

multimodal
by Nex Agi · 262,144 ctx

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual fee...

NVIDIA: Switchyard

text
by Nvidia · 1,000,000 ctx

Switchyard is an open-source model router that switches between multiple models to optimize the cost of requests. By default it will use ...

OpenAI: GPT-6.1 Sol

multimodal
by OpenAI · 1,050,000 ctx

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agent...

OpenAI: GPT-6.1 Sol Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-6.1 Sol Pro is the same underlying model as [GPT-6.1 Sol](https://openrouter.ai/openai/gpt-6.1-sol), served with `reasoning.mode` set...

OpenAI: GPT-6 Astra

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep rese...

OpenAI: GPT-6 Astra Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set...

OpenAI: GPT-6 Luna

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and laten...

OpenAI: GPT-6 Luna Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to...

OpenAI: GPT-6 Sol

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT...

OpenAI: GPT-6 Sol Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `p...

?

OpenAI GPT Astra Latest

multimodal
by ~openai · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT Astra family.

?

OpenAI GPT Luna Latest

multimodal
by ~openai · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT Luna family.

?

OpenAI GPT Sol Latest

multimodal
by ~openai · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT Sol family.

?

OpenAI GPT Terra Latest

multimodal
by ~openai · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT Terra family.

?

Pareto

multimodal
by Unbiased · 262,144 ctx

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance acr...

?

Pareto 26.10 Preview

multimodal
by Unbiased · 1,048,576 ctx

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance acr...

?

Perceptron: Perceptron Mk1.5

multimodal
by Perceptron · 36,864 ctx

Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answer...

?

PrismML: Ternary Bonsai 2 27B

multimodal
by Prism ML · 262,144 ctx

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and...

Qwen: Qwen3.8 Max Prime

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price poin...

Qwen: Qwen3.8 Omni Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audi...

?

Sakana: Fugu Max

multimodal
by Sakana · 1,000,000 ctx

Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent o...

?

Sakana: Fugu Ultra v2

multimodal
by Sakana · 1,000,000 ctx

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-...

?

Space Bunny Alpha

multimodal
by Stealth · 1,000,000 ctx

Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support...

SpaceXAI: Grok 4.7

multimodal
by xAI · 500,000 ctx

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at lo...

?

TypeSafe: Jev Router

multimodal
by Typesafe · 1,000,000 ctx

Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. It runs on [Jev](https://openr...

?

Union Alpha

multimodal
by Stealth · 262,144 ctx

Union Alpha is a multimodal model built for research, coding, and agentic workflows, while delivering frontier-level performance across a...

Upstage: Solar Mini 4

text
by Upstage · 524,288 ctx

Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K...

Xiaomi: MiMo-V2.6-Flash

multimodal
by Xiaomi · 1,048,576 ctx

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parame...

Xiaomi: MiMo-V2.6-Pro

multimodal
by Xiaomi · 1,048,576 ctx

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ce...

Xiaomi: MiMo-V2.6-Pro-UltraSpeed

multimodal
by Xiaomi · 1,048,576 ctx

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6...

Z.ai: GLM 5.3 FlashX

multimodal
by Zhipu AI · 1,048,576 ctx

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 toke...

Z.ai: GLM 5.3 Prime

text
by Zhipu AI · 1,000,000 ctx

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput...

OpenAI: GPT-3.5 Turbo

text
by OpenAI · 16,385 ctx

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and tradition...

OpenAI: GPT-3.5 Turbo (older v0613)

text
by OpenAI · 4,095 ctx

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and tradition...

OpenAI: GPT-3.5 Turbo 16k

text
by OpenAI · 16,385 ctx

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single reque...

OpenAI: GPT-3.5 Turbo Instruct

text
by OpenAI · 4,095 ctx

This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Se...

OpenAI: GPT-4

text
by OpenAI · 8,191 ctx

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy tha...

OpenAI: GPT-4 (older v0314)

text
by OpenAI · 8,191 ctx

GPT-4-0314 is the first version of GPT-4 released, with a context length of 8,192 tokens, and was supported until June 14. Training data:...

OpenAI: GPT-4 Turbo (older v1106)

text
by OpenAI · 128,000 ctx

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to ...

OpenAI: GPT-4.1

multimodal
by OpenAI · 1,047,576 ctx

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-contex...

OpenAI: GPT-4.1 Mini

multimodal
by OpenAI · 1,047,576 ctx

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 ...

OpenAI: GPT-4.1 Nano

multimodal
by OpenAI · 1,047,576 ctx

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performa...

OpenAI: GPT-4o (2024-05-13)

multimodal
by OpenAI · 128,000 ctx

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligen...

OpenAI: GPT-4o (2024-08-06)

multimodal
by OpenAI · 128,000 ctx

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respo...

OpenAI: GPT-4o (2024-11-20)

multimodal
by OpenAI · 128,000 ctx

The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improv...

OpenAI: GPT-4o-mini (2024-07-18)

multimodal
by OpenAI · 128,000 ctx

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. ...

OpenAI: GPT-5 Chat

multimodal
by OpenAI · 128,000 ctx

GPT-5 Chat is designed for advanced, natural, multimodal, and context-aware conversations for enterprise applications.

OpenAI: GPT-5 Codex

multimodal
by OpenAI · 400,000 ctx

GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactiv...

OpenAI: GPT-5 Image

multimodal
by OpenAI · 400,000 ctx

[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It o...

OpenAI: GPT-5 Image Mini

multimodal
by OpenAI · 400,000 ctx

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with...

OpenAI: GPT-5 Mini

multimodal
by OpenAI · 400,000 ctx

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following a...

OpenAI: GPT-5 Nano

multimodal
by OpenAI · 400,000 ctx

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low late...

OpenAI: GPT-5 Pro

multimodal
by OpenAI · 400,000 ctx

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized f...

OpenAI: GPT-5.1

multimodal
by OpenAI · 400,000 ctx

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adheren...

OpenAI: GPT-5.1 Chat

multimodal
by OpenAI · 128,000 ctx

GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong genera...

OpenAI: GPT-5.1-Codex

multimodal
by OpenAI · 400,000 ctx

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both intera...

OpenAI: GPT-5.1-Codex-Max

multimodal
by OpenAI · 400,000 ctx

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is base...

OpenAI: GPT-5.1-Codex-Mini

multimodal
by OpenAI · 400,000 ctx

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

OpenAI: GPT-5.2

multimodal
by OpenAI · 400,000 ctx

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1...

OpenAI: GPT-5.2 Chat

multimodal
by OpenAI · 128,000 ctx

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong gener...

OpenAI: GPT-5.2 Pro

multimodal
by OpenAI · 400,000 ctx

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. I...

OpenAI: GPT-5.2-Codex

multimodal
by OpenAI · 400,000 ctx

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both in...

OpenAI: GPT-5.3 Chat

multimodal
by OpenAI · 128,000 ctx

GPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful...

OpenAI: GPT-5.3-Codex

multimodal
by OpenAI · 400,000 ctx

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex wi...

OpenAI: GPT-5.4

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window ...

OpenAI: GPT-5.4 Image 2

multimodal
by OpenAI · 272,000 ctx

[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabiliti...

OpenAI: GPT-5.4 Mini

multimodal
by OpenAI · 400,000 ctx

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It suppor...

OpenAI: GPT-5.4 Nano

multimodal
by OpenAI · 400,000 ctx

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks...

OpenAI: GPT-5.4 Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex,...

OpenAI: GPT-5.5

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher relia...

OpenAI: GPT-5.5 Pro

multimodal
by OpenAI · 1,050,000 ctx

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a ...

OpenAI: gpt-oss-safeguard-20b

20B
by OpenAI · 131,072 ctx

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts ...

OpenAI: o1

multimodal
by OpenAI · 200,000 ctx

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is t...

OpenAI: o1-pro

multimodal
by OpenAI · 200,000 ctx

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro mod...

OpenAI: o3

multimodal
by OpenAI · 200,000 ctx

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It ...

OpenAI: o3 Deep Research

multimodal
by OpenAI · 200,000 ctx

o3-deep-research is OpenAI's advanced model for deep research, designed to tackle complex, multi-step research tasks. Note: This model a...

OpenAI: o3 Mini

text
by OpenAI · 200,000 ctx

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

OpenAI: o3 Mini High

text
by OpenAI · 200,000 ctx

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient langua...

OpenAI: o3 Pro

multimodal
by OpenAI · 200,000 ctx

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro mode...

OpenAI: o4 Mini

multimodal
by OpenAI · 200,000 ctx

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multim...

OpenAI: o4 Mini Deep Research

multimodal
by OpenAI · 200,000 ctx

o4-mini-deep-research is OpenAI's faster, more affordable deep research model—ideal for tackling complex, multi-step research tasks. Not...

OpenAI: o4 Mini High

multimodal
by OpenAI · 200,000 ctx

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reason...

?

Owl Alpha

text
by Openrouter · 1,048,756 ctx

Owl Alpha is a high-performance foundation model designed for agentic workloads. Natively supports tool use, and long-context tasks, with...

?

Pareto Code Router

text
by Openrouter · 2,000,000 ctx

The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) c...

Perplexity: Sonar

multimodal
by Perplexity · 127,072 ctx

Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed ...

Perplexity: Sonar Deep Research

text
by Perplexity · 128,000 ctx

Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It aut...

Perplexity: Sonar Pro

multimodal
by Perplexity · 200,000 ctx

Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing...

Perplexity: Sonar Pro Search

multimodal
by Perplexity · 200,000 ctx

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is d...

Perplexity: Sonar Reasoning Pro

multimodal
by Perplexity · 128,000 ctx

Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing...

xAI: Grok 4.20

multimodal
by xAI · 2,000,000 ctx

Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest halluci...

xAI: Grok 4.20 Multi-Agent

multimodal
by xAI · 2,000,000 ctx

Grok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in paral...

xAI: Grok 4.3

multimodal
by xAI · 1,000,000 ctx

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instructi...

xAI: Grok Build 0.1

multimodal
by xAI · 256,000 ctx

Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inp...

Claude Opus 4.7

text
by Anthropic · Claude · 200,000 ctx

Frontier reasoning and long-form coding from Anthropic.

Claude Sonnet 4.6

text
by Anthropic · Claude · 200,000 ctx

Best price-performance from Anthropic. Default for production agents.

Claude Haiku 4.5

text
by Anthropic · Claude · 200,000 ctx

Fast, cheap Claude variant for high-throughput inference.

GPT-5

text
by OpenAI · GPT · 256,000 ctx

OpenAI's frontier multimodal reasoning model.

Gemini 2.5 Pro

multimodal
by Google DeepMind · Gemini · 1,000,000 ctx

Google's frontier reasoning model with native 1M-token context.

Grok 3

multimodal
by xAI · Grok · 1,000,000 ctx

xAI's frontier model with built-in DeepSearch + real-time X integration.

Claude 3.5 Haiku

text
by Anthropic · Claude · 200,000 ctx

Fast/cheap Claude 3.5 variant — production fallback for Haiku 4.5.

GPT-4o Mini

multimodal
by OpenAI · GPT · 128,000 ctx

Cheap multimodal default — replaced GPT-3.5 Turbo for low-cost workloads.

Claude 3.5 Sonnet

text
by Anthropic · Claude · 200,000 ctx

Anthropic's 3.5 generation — still in active production.

Gemini 1.5 Flash

multimodal
by Google DeepMind · Gemini · 1,000,000 ctx

Cheap fast Gemini — production default before 2.0/2.5 Flash.

GPT-4o

multimodal
by OpenAI · GPT · 128,000 ctx

OpenAI's multimodal model — text, vision, audio in one.

GPT-4 Turbo

text
by OpenAI · GPT · 128,000 ctx

OpenAI's pre-GPT-5 flagship — still extensively deployed.

Gemini 1.5 Pro

multimodal
by Google DeepMind · Gemini · 1,000,000 ctx

Google's pre-2.5 frontier — 2M context launched here.

?

Z.ai: GLM Flash Latest

multimodal
by ~z Ai · 1,310,720 ctx

This model always redirects to the latest model in the GLM Flash family.

?

Z.ai: GLM Latest

text
by ~z Ai · 1,048,576 ctx

This model always redirects to the latest GLM model from Z.ai.

?

DeepSeek V4 Flash Latest

text
by ~deepseek · 1,048,576 ctx

This model always redirects to the latest model in the DeepSeek V4 Flash family.

Claude Opus 5 (Fast)

multimodal
by Anthropic · 1,000,000 ctx

Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regul...

?

xAI: Grok Latest

multimodal
by ~x Ai · 500,000 ctx

This model always redirects to the latest Grok model from xAI.

?

Anthropic: Claude Fable Latest

multimodal
by ~anthropic · 1,000,000 ctx

This model always redirects to the latest model in the Claude Fable family.

Anthropic: Claude Opus 4.8 (Fast)

multimodal
by Anthropic · 1,000,000 ctx

Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to r...

Anthropic Claude Haiku Latest

multimodal
by Anthropic · 200,000 ctx

This model always redirects to the latest model in the Anthropic Claude Haiku family.

Anthropic Claude Sonnet Latest

multimodal
by Anthropic · 1,000,000 ctx

This model always redirects to the latest model in the Anthropic Claude Sonnet family.

Anthropic: Claude Opus 4.6 (Fast)

multimodal
by Anthropic · 1,000,000 ctx

Fast-mode variant of [Opus 4.6](/anthropic/claude-opus-4.6) - identical capabilities with higher output speed at premium 6x pricing. Lea...

Anthropic: Claude Opus 4.7 (Fast)

multimodal
by Anthropic · 1,000,000 ctx

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Lea...

Anthropic: Claude Opus Latest

multimodal
by Anthropic · 1,000,000 ctx

This model always redirects to the latest model in the Claude Opus family.

?

Google Gemini Flash Latest

multimodal
by ~google · 1,048,576 ctx

This model always redirects to the latest model in the Google Gemini Flash family.

?

Google Gemini Pro Latest

multimodal
by ~google · 1,048,576 ctx

This model always redirects to the latest model in the Google Gemini Pro family.

Google: Gemini 2.5 Flash Lite Preview 09-2025

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It ...

Google: Gemini 2.5 Pro Preview 05-06

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It emplo...

Google: Gemini 2.5 Pro Preview 06-05

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It emplo...

Google: Gemini 3 Flash Preview

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance....

Google: Gemini 3.1 Flash Lite Preview

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite...

Google: Gemini 3.1 Pro Preview

multimodal
by Google DeepMind · 1,048,576 ctx

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic relia...

Google: Gemini 3.1 Pro Preview Custom Tools

multimodal
by Google DeepMind · 1,048,756 ctx

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a gener...

Google: Lyria 3 Clip Preview

multimodal
by Google DeepMind · 1,048,576 ctx

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemin...

Google: Lyria 3 Pro Preview

multimodal
by Google DeepMind · 1,048,576 ctx

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. ...

Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)

multimodal
by Google DeepMind · 131,072 ctx

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, deliverin...

Google: Nano Banana Pro (Gemini 3 Pro Image Preview)

multimodal
by Google DeepMind · 65,536 ctx

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana ...

OpenAI GPT Latest

multimodal
by OpenAI · 1,050,000 ctx

This model always redirects to the latest model in the OpenAI GPT family.

OpenAI GPT Mini Latest

multimodal
by OpenAI · 400,000 ctx

This model always redirects to the latest model in the OpenAI GPT Mini family.

OpenAI: GPT Chat Latest

multimodal
by OpenAI · 400,000 ctx

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. ...

OpenAI: GPT-4 Turbo Preview

text
by OpenAI · 128,000 ctx

The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Traini...

OpenAI: GPT-4o Search Preview

text
by OpenAI · 128,000 ctx

GPT-4o Search Previewis a specialized model for web search in Chat Completions. It is trained to understand and execute web search queries.

OpenAI: GPT-4o-mini Search Preview

text
by OpenAI · 128,000 ctx

GPT-4o mini Search Preview is a specialized model for web search in Chat Completions. It is trained to understand and execute web search ...