AI models

Every way
to use the major models.

Closed models like Claude and GPT — link to the cheapest API provider. Open-weights like Llama, Kimi, DeepSeek — choose hosted inference or self-host on rented GPUs.

507 tracked · 507 open weights · 0 closed APIs · cheapest input $0.01/M
Quality × Price

Find the sweet spot.

Higher = stronger benchmark composite · further left = cheaper input

Loading...

507 models match — reset filters

Open-weights models.

Run yourself on cheap GPUs, or use a hosted-inference provider.

Meta: Muse Spark 1.3

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to k...

?

GLM 5.3 Flash

multimodal
by zai-org · 1,048,576 ctx

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. It incorpo...

?

GLM-5.3

text
by zai-org · 1,048,576 ctx

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex co...

Qwen: Qwen3.8 Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, ...

Z.ai: GLM 5.3 Flash

multimodal
by Zhipu AI · 1,310,720 ctx

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse a...

granite-4.2-30b

30B
by IBM Research · 131,072 ctx

Granite-4.2-30B is the flagship reasoning model in the Granite 4.2 family. It delivers the strongest performance across reasoning-intensi...

granite-4.2-3b

3B
by IBM Research · 131,072 ctx

Granite-4.2-3B is the compact reasoning model in the Granite 4.2 family. Despite its small parameter count, it delivers strong performanc...

granite-4.2-8b

9B
by IBM Research · 131,072 ctx

Granite-4.2-8B is the mid-size reasoning model in the Granite 4.2 family. It delivers strong performance on reasoning-intensive tasks by ...

Tencent: Hy-MT2-7B

7B
by Tencent · 8,192 ctx

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language ...

Mistral: Ministral 8B

8B
by Mistral AI · 128,000 ctx

Ministral 8B is an 8B parameter model featuring a unique interleaved sliding-window attention pattern for faster, memory-efficient infere...

Tencent: Hy-MT2-1.8B

8B
by Tencent · 8,192 ctx

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minori...

Tencent: Hy-MT2-30B-A3B

30B
by Tencent · 8,192 ctx

Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and m...

glm-5.3

text
by Zhipu AI · 1,048,576 ctx

Qwen: Qwen3.8 27B

27B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimoda...

Qwen3.8-2.4T-A95B

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.8-2.4T-A95B is Alibaba's most capable Qwen model to date, a 2.4T-parameter sparse MoE with ~95B active parameters. It is built for ...

?

ByteDance Seed: Seed 2.1 Turbo

multimodal
by Bytedance Seed · 262,144 ctx

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end softwar...

DeepSeek: DeepSeek V4 Pro 0813

text
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Qwen3.8-2.4T-A95B

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.8-2.4T-A95B is Alibaba’s most capable Qwen model to date, a 2.4T-parameter sparse MoE with ~95B active parameters. It is built for ...

Qwen: Qwen3.8 2.4T A95B

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-...

NVIDIA-Nemotron-3.5-Lightning

text
by Nvidia · 28,672 ctx

NVIDIA Nemotron 3.5 Lightning is NVIDIA's fastest open model for always-on agents and high-volume specialized tasks. It delivers a substa...

NVIDIA: Nemotron 3.5 Lightning

text
by Nvidia · 262,144 ctx

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited f...

Nemotron Lightning 3.5 30B A3B

30B
by Nvidia · 262,144 ctx

Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on...

Meta: Muse Glimmer 30B

30B
by Meta AI · 131,072 ctx

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for a...

Upstage: Solar Pro 4

text
by Upstage · 524,288 ctx

Solar Pro 4 is a large language model from Upstage. It is suited for agentic workflows, office productivity, document-intensive work, and...

?

Ling-3.0-flash

text
by Inclusionai · 131,072 ctx

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The mode...

Meta: Muse Spark 1.2

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documen...

Qwen: Qwen3.8 Max

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a mul...

DeepSeek: DeepSeek V4 Flash 0731

text
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-tra...

Qwen: Qwen3.7 Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer ...

Meta: Muse Spark 1.1

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documen...

?

MoonshotAI: Kimi K3

multimodal
by Moonshot AI · 1,048,576 ctx

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and...

?

Ternary Bonsai 27B

27B
by Prism ML · 262,144 ctx

Kwaipilot: KAT-Coder-Air V2.5

text
by Kwaipilot · 256,000 ctx

Kwaipilot: KAT-Coder-Pro V2.5

text
by Kwaipilot · 256,000 ctx

Gemma 4 12B It

12B
by Google DeepMind · 262,144 ctx
?

LFM2.5-8B-A1B

8B
by LiquidAI · 128,000 ctx
?

AionLabs: Aion-3.0

text
by Aion Labs · 131,072 ctx

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative g...

?

AionLabs: Aion-3.0-Mini

text
by Aion Labs · 131,072 ctx

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a colla...

?

Nex AGI: Nex-N2-Mini

multimodal
by Nex Agi · 262,144 ctx

Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text an...

Tencent: Hy3

text
by Tencent · 262,144 ctx

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic w...

?

Ornith-1.0-35B

35B
by Deepreinforce Ai · 262,144 ctx

Ornith-1.0-35B is DeepReinforce's open (MIT-licensed) agentic-coding model: an RL post-train of Qwen3.5-35B-A3B, a 35B-total / ~3B-active...

?

Nex AGI: Nex-N2-Pro

multimodal
by Nex Agi · 262,144 ctx

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 archit...

?

GLM 5.2

text
by zai-org · 1,048,576 ctx

GLM-5.2 introduces a robust 1M-token context and advanced, multi-effort coding capabilities to significantly enhance performance on long-...

Z.ai: GLM 5.2

text
by Zhipu AI · 1,048,576 ctx

GLM-5.2 is Z.ai’s flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-le...

?

Kimi 2.7 Code

multimodal
by Moonshot AI · 262,144 ctx

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding ta...

?

MoonshotAI: Kimi K2.7 Code

multimodal
by Moonshot AI · 262,144 ctx

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reli...

NVIDIA-Nemotron-3-Ultra-550B-A55B

550B
by Nvidia · 262,144 ctx

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It deli...

NVIDIA: Nemotron 3 Ultra

550B
by Nvidia · 1,000,000 ctx

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (...

Nemotron-Content-Safety-3.5

multimodal
by Nvidia · 131,072 ctx

Nemotron Content Safety 3.5 is a multimodal safety classifier developed by NVIDIA. A compact safety model that handles text, images, and...

Qwen: Qwen3.7 Plus

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the se...

MiniMax: MiniMax M3

multimodal
by MiniMax · 1,048,576 ctx

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context ...

StepFun: Step 3.7 Flash

multimodal
by Stepfun · 256,000 ctx

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with ...

?

AionLabs: Aion-1.0

text
by Aion Labs · 131,072 ctx

Aion-1.0 is a multi-model system designed for high performance across various tasks, including reasoning and coding. It is built on DeepS...

?

AionLabs: Aion-1.0-Mini

text
by Aion Labs · 131,072 ctx

Aion-1.0-Mini 32B parameter model is a distilled version of the DeepSeek-R1 model, designed for strong performance in reasoning domains s...

?

AionLabs: Aion-2.0

text
by Aion Labs · 131,072 ctx

Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing ten...

?

AionLabs: Aion-RP 1.0 (8B)

8B
by Aion Labs · 32,768 ctx

Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant ...

?

AlfredPros: CodeLLaMa 7B Instruct Solidity

7B
by Alfredpros · 4,096 ctx

A finetuned 7 billion parameters Code LLaMA - Instruct model to generate Solidity smart contract using 4-bit QLoRA finetuning provided by...

AllenAI: Olmo 3 32B Think

32B
by Allen Institute for AI (AI2) · 65,536 ctx

Olmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instruc...

Arcee AI: Coder Large

text
by Arcee Ai · 32,768 ctx

Coder‑Large is a 32 B‑parameter offspring of Qwen 2.5‑Instruct that has been further trained on permissively‑licensed GitHub, CodeSearchN...

Arcee AI: Maestro Reasoning

text
by Arcee Ai · 131,072 ctx

Maestro Reasoning is Arcee's flagship analysis model: a 32 B‑parameter derivative of Qwen 2.5‑32 B tuned with DPO and chain‑of‑thought RL...

Arcee AI: Spotlight

1B
by Arcee Ai · 131,072 ctx

Spotlight is a 7‑billion‑parameter vision‑language model derived from Qwen 2.5‑VL and fine‑tuned by Arcee AI for tight image‑text groundi...

Arcee AI: Trinity Large Thinking

399B
by Arcee Ai · 262,144 ctx

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, ag...

Arcee AI: Trinity Mini

text
by Arcee Ai · 131,072 ctx

Trinity Mini is a 26B-parameter (3B active) sparse mixture-of-experts language model featuring 128 experts with 8 active per token. Engin...

Arcee AI: Virtuoso Large

text
by Arcee Ai · 131,072 ctx

Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters, tuned to tackle cross‑domain reasoning, creative writing and e...

Arize AI Qwen 2 1.5B Instruct

5B
by Togethercomputer · 32,768 ctx

Baidu: ERNIE 4.5 21B A3B

21B
by Baidu · 131,072 ctx

A sophisticated text-based Mixture-of-Experts (MoE) model featuring 21B total parameters with 3B activated per token, delivering exceptio...

Baidu: ERNIE 4.5 21B A3B Thinking

21B
by Baidu · 131,072 ctx

ERNIE-4.5-21B-A3B-Thinking is Baidu's upgraded lightweight MoE model, refined to boost reasoning depth and quality for top-tier performan...

Baidu: ERNIE 4.5 300B A47B

300B
by Baidu · 131,072 ctx

ERNIE-4.5-300B-A47B is a 300B parameter Mixture-of-Experts (MoE) language model developed by Baidu as part of the ERNIE 4.5 series. It ac...

Baidu: ERNIE 4.5 VL 28B A3B

28B
by Baidu · 131,072 ctx

A powerful multimodal Mixture-of-Experts chat model featuring 28B total parameters with 3B activated per token, delivering exceptional te...

Baidu: ERNIE 4.5 VL 424B A47B

424B
by Baidu · 131,072 ctx

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with...

Baidu: Qianfan-OCR-Fast

multimodal
by Baidu · 65,536 ctx

Qianfan-OCR-Fast is a domain-specific multimodal large model purpose-built for OCR. By leveraging specialized OCR training data while pre...

?

ByteDance Seed: Seed 1.6

200B
by Bytedance Seed · 262,144 ctx

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinki...

?

ByteDance Seed: Seed 1.6 Flash

multimodal
by Bytedance Seed · 262,144 ctx

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It featu...

?

ByteDance Seed: Seed-2.0-Lite

32B
by Bytedance Seed · 262,144 ctx

Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering n...

?

ByteDance Seed: Seed-2.0-Mini

multimodal
by Bytedance Seed · 262,144 ctx

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference...

?

ByteDance: UI-TARS 7B

7B
by Bytedance · 128,000 ctx

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobil...

Deep Cogito: Cogito v2.1 671B

671B
by Deepcogito · 128,000 ctx

Cogito v2.1 671B MoE represents one of the strongest open models globally, matching performance of frontier closed and open models. This ...

DeepSeek-V3-0324

671B
by DeepSeek · 163,840 ctx

DeepSeek-V3-0324, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token, an impro...

DeepSeek-V3.1

text
by DeepSeek · 163,840 ctx

DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase l...

DeepSeek: DeepSeek V3

689B
by DeepSeek · 163,840 ctx

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous vers...

DeepSeek: DeepSeek V3 0324

689B
by DeepSeek · 163,840 ctx

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team...

DeepSeek: DeepSeek V3.1

689B
by DeepSeek · 163,840 ctx

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prom...

DeepSeek: DeepSeek V3.1 Terminus

689B
by DeepSeek · 163,840 ctx

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities whi...

DeepSeek: DeepSeek V3.2

689B
by DeepSeek · 131,072 ctx

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use pe...

DeepSeek: DeepSeek V3.2 Speciale

text
by DeepSeek · 163,840 ctx

DeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance. It builds on D...

DeepSeek: DeepSeek V4 Flash

292B
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated paramete...

DeepSeek: DeepSeek V4 Pro

1602B
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporti...

DeepSeek: R1 0528

671B
by DeepSeek · 163,840 ctx

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced an...

Deepseek Coder 33B Instruct

33B
by DeepSeek · 16,384 ctx

Deepseek OCR 2

text
by DeepSeek · 8,192 ctx

Devstral Small 2505

text
by Mistral AI · 131,072 ctx

EssentialAI Rnj-1 Instruct

text
by Essentialai · 32,768 ctx

EssentialAI: Rnj 1 Instruct

text
by Essentialai · 32,768 ctx

Rnj-1 is an 8B-parameter, dense, open-weight model family developed by Essential AI and trained from scratch with a focus on programming,...

Facebook CWM

text
by Meta AI · 131,072 ctx
?

GLM 5.1

text
by zai-org · 202,752 ctx

GLM-5.1 is Z.ai's next-generation flagship model built for agentic engineering, with stronger coding capabilities and sustained performan...

GLM OCR

text
by Zhipu AI · 131,072 ctx

GLM-4.5-Flash

text
by Zhipu AI · GLM

Lowest-latency, lowest-cost variant of GLM-4.5 on z.ai.

GLM-4.6

355B
by Zhipu AI · GLM · 202,752 ctx

Incremental upgrade on GLM-4.5 — improved reasoning, same context window.

GLM-4.7

9B
by Zhipu AI · GLM · 202,752 ctx

Mid-generation GLM 4.7 released between GLM-4.6 and GLM-5.

GLM-5

355B
by Zhipu AI · GLM · 202,752 ctx

Zhipu's GLM 5 generation — closed flagship between GLM-4.7 and GLM-5.1.

GLM-5 Turbo

text
by Zhipu AI · GLM · 202,752 ctx

Faster, cheaper sibling of GLM-5 on z.ai.

GLM-5.1

32B
by Zhipu AI · GLM · 202,752 ctx

Zhipu's GLM 5.1 series — successor to GLM-5 on z.ai's API.

Gemma 2 9B It

9B
by Google DeepMind · 8,192 ctx

Gemma 2B It

2B
by Google DeepMind · 8,192 ctx

Gemma 3 1B Pt

1B
by Google DeepMind · 32,768 ctx

Gemma 3 1b it

1B
by Google DeepMind · 32,768 ctx

Gemma 3 270M It

0B
by Google DeepMind · 32,768 ctx

Gemma 3 27B It

27B
by Google DeepMind · 65,536 ctx

Gemma 3 27B Pt

27B
by Google DeepMind

Gemma 3 4b it

4B
by Google DeepMind · 65,536 ctx

Gemma 4 E2B-it

text
by Google DeepMind · 131,072 ctx

Gemma 4 E4B-it

text
by Google DeepMind · 131,072 ctx
?

Hermes-3-Llama-3.1-70B

70B
by Nous Research · 131,072 ctx

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better rolepl...

?

Holo3 35B A3b

35B
by Hcompany · 262,144 ctx

IBM: Granite 4.0 Micro

3B
by IBM Research · 131,000 ctx

Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by ...

IBM: Granite 4.1 8B

8B
by IBM Research · 131,072 ctx

Granite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-t...

?

Inception: Mercury 2

text
by Inception · 128,000 ctx

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Me...

?

Kimi K2.5

multimodal
by Moonshot AI · 262,144 ctx

Kimi K2.5 is Moonshot AI's flagship agentic model and a new SOTA open model. It unifies vision and text, thinking and non-thinking modes,...

?

Kimi K2.6

multimodal
by Moonshot AI · 262,144 ctx

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven d...

Kwaipilot: KAT-Coder-Pro V2

text
by Kwaipilot · 256,000 ctx

KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engin...

?

L3.1-70B-Euryale-v2.2

70B
by Sao10k · 131,072 ctx

Euryale 3.1 - 70B v2.2 is a model focused on creative roleplay from Sao10k

LFM2-24B-A2B

24B
by Togethercomputer · 32,768 ctx

LiquidAI: LFM2-24B-A2B

24B
by Liquid · 128,000 ctx

LFM2-24B-A2B is the largest model in the LFM2 family of hybrid architectures designed for efficient on-device deployment. Built as a 24B ...

Llama 3.1 70B

70B
by Meta AI · 131,072 ctx

Llama 3.1 Nemotron 70B Instruct HF

70B
by Nvidia · 32,768 ctx

Llama 4 Scout (17Bx16E)

17B
by Meta AI · 262,144 ctx

Llama 4 Scout Instruct (17Bx16E)

17B
by Meta AI · 1,048,576 ctx

Llama Guard 3 8B

8B
by Meta AI · 131,072 ctx

Llama Guard 3 is a Llama-3.1-8B pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be u...

Llama-3.2-11B-Vision-Instruct

11B
by Meta AI · 131,072 ctx

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It exc...

Magistral Small 2506

text
by Mistral AI · 40,960 ctx
?

Magnum v4 72B

72B
by Anthracite Org · 32,768 ctx

This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anth...

?

Mancer: Weaver (alpha)

text
by Mancer · 8,000 ctx

An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrativ...

Medgemma 27B Text It

27B
by Google DeepMind · 131,072 ctx

DeepSeek: DeepSeek V4.1 Flash

multimodal
by DeepSeek · 1,048,576 ctx

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED)...

DeepSeek V4.1 Flash

multimodal
by DeepSeek · 1,048,576 ctx

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text at up ...

?

Ember-1

multimodal
by fireworks · 1,048,576 ctx

Ember-1 is a specialized model from Fireworks. Built on Kimi K3, it produces shorter reasoning traces, using approximately 40% fewer toke...

?

Inception: Mercury 2.5

text
by Inception · 260,000 ctx

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, ...

?

Ling 3.0 Flash Fin

text
by Inclusionai · 262,144 ctx

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters ou...

?

Ling-3.0-flash-VL

125B
by Inclusionai · 131,072 ctx

The multimodal version built on Ling-3.0-flash — 124B total / ~5.5B active per token, with native text, image, and video understanding. I...

Meta Llama 3 70B Instruct Turbo

70B
by Meta AI · 8,192 ctx

NVIDIA: Nemotron 3.5 Content Safety

multimodal
by Nvidia · 131,072 ctx

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. I...

Qwen: Qwen3.8 Max (0902)

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model ...

?

Tev1 4B Experimental

text
by Together AI · 32,768 ctx

Meta Llama 3 8B Instruct

8B
by Meta AI · 8,192 ctx

Meta Llama 3 8B Instruct Lite

8B
by Meta AI · 8,192 ctx

Meta Llama 3 8B Instruct Reference

8B
by Meta AI · 8,192 ctx

Meta Llama 3.1 405B Instruct

405B
by Meta AI · 4,096 ctx

Meta Llama 3.1 70B Instruct Turbo

70B
by Meta AI · 131,072 ctx

Meta Llama 3.1 8B

8B
by Meta AI · 16,384 ctx

Meta Llama 3.1 8B Instruct Turbo

8B
by Meta AI · 131,072 ctx

Meta Llama 3.2 1B Instruct

1B
by Meta AI · 131,072 ctx

Meta Llama 3.2 3B Instruct

3B
by Meta AI · 131,072 ctx

Meta Llama 3.3 70B Instruct Turbo

70B
by Meta AI · 131,072 ctx

Meta-Llama-3.1-70B-Instruct

70B
by Meta AI · 131,072 ctx

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned g...

Meta-Llama-3.1-8B-Instruct

8B
by Meta AI · 131,072 ctx

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned g...

Meta: Llama 3 70B Instruct

70B
by Meta AI · 8,192 ctx

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 70B instruct-tuned version was optimized for high...

Meta: Llama 3 8B Instruct

8B
by Meta AI · 8,192 ctx

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 8B instruct-tuned version was optimized for high ...

Meta: Llama 4 Maverick

402B
by Meta AI · 1,048,576 ctx

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architec...

Meta: Llama 4 Scout

109B
by Meta AI · 10,000,000 ctx

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of ...

Meta: Llama Guard 4 12B

12B
by Meta AI · 163,840 ctx

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous v...

Microsoft: Phi 4 Mini Instruct

4B
by Microsoft · 131,072 ctx

Phi-4-mini-instruct is a lightweight open model built upon synthetic data and filtered publicly available websites - with a focus on high...

MiniMax M2.7

456B
by MiniMax · 196,608 ctx

Mixture-of-Experts language model. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks...

MiniMax-M2.5

text
by MiniMax · 196,608 ctx

MiniMax M2.5 is built for state-of-the-art coding, agentic tool use, search, and office work, extensively trained with reinforcement lear...

MiniMax: MiniMax M1

456B
by MiniMax · 1,000,000 ctx

MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybr...

MiniMax: MiniMax M2

456B
by MiniMax · 204,800 ctx

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion acti...

MiniMax: MiniMax M2-her

text
by MiniMax · 65,536 ctx

MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn co...

MiniMax: MiniMax M2.1

230B
by MiniMax · 204,800 ctx

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application deve...

MiniMax: MiniMax M2.5

230B
by MiniMax · 204,800 ctx

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digita...

MiniMax: MiniMax M2.7

230B
by MiniMax · 204,800 ctx

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built...

MiniMax: MiniMax-01

457B
by MiniMax · 1,000,192 ctx

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, wi...

Minimax M1 40K

text
by MiniMax · 1,048,576 ctx

Minimax M1 80K

text
by MiniMax · 1,048,576 ctx

Ministral 3 14B Instruct 2512

14B
by Mistral AI · 262,144 ctx

Mistral (7B) Instruct v0.3

7B
by Mistral AI · 32,768 ctx

Mistral Large

text
by Mistral AI · 128,000 ctx

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excel...

Mistral Large 2407

text
by Mistral AI · 131,072 ctx

This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels ...

Mistral-Nemo-Instruct-2407

text
by Mistral AI · 131,072 ctx

12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.

Mistral-Small-3.2-24B-Instruct-2506

24B
by Mistral AI · 128,000 ctx

Mistral-Small-3.2-24B-Instruct is a drop-in upgrade over the 3.1 release, with markedly better instruction following, roughly half the in...

Mistral: Codestral 2508

text
by Mistral AI · 256,000 ctx

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks su...

Mistral: Devstral 2 2512

128B
by Mistral AI · 262,144 ctx

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer...

Mistral: Devstral Medium

24B
by Mistral AI · 131,072 ctx

Devstral Medium is a high-performance code generation and agentic reasoning model developed jointly by Mistral AI and All Hands AI. Posit...

Mistral: Devstral Small 1.1

24B
by Mistral AI · 131,072 ctx

Devstral Small 1.1 is a 24B parameter open-weight language model for software engineering agents, developed by Mistral AI in collaboratio...

Mistral: Ministral 3 14B 2512

14B
by Mistral AI · 262,144 ctx

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistra...

Mistral: Ministral 3 3B 2512

3B
by Mistral AI · 131,072 ctx

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

Mistral: Ministral 3 8B 2512

8B
by Mistral AI · 262,144 ctx

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

Mistral: Mistral 7B Instruct v0.1

7B
by Mistral AI · 4,096 ctx

A 7.3B parameter model that outperforms Llama 2 13B on all benchmarks, with optimizations for speed and context length.

Mistral: Mistral Large 3 2512

multimodal
by Mistral AI · 262,144 ctx

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active paramete...

Mistral: Mistral Medium 3

multimodal
by Mistral AI · 131,072 ctx

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly r...

Mistral: Mistral Medium 3.1

multimodal
by Mistral AI · 131,072 ctx

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to del...

Mistral: Mistral Medium 3.5

multimodal
by Mistral AI · 262,144 ctx

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and i...

Mistral: Mistral Small 3

24B
by Mistral AI · 32,768 ctx

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache...

Mistral: Mistral Small 3.1 24B

24B
by Mistral AI · 128,000 ctx

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal...

Mistral: Mistral Small 3.2 24B

24B
by Mistral AI · 128,000 ctx

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduct...

Mistral: Mistral Small 4

121B
by Mistral AI · 262,144 ctx

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into ...

Mistral: Pixtral Large 2411

multimodal
by Mistral AI · 131,072 ctx

Pixtral Large is a 124B parameter, open-weight, multimodal model built on top of [Mistral Large 2](/mistralai/mistral-large-2411). The mo...

Mistral: Saba

text
by Mistral AI · 32,768 ctx

Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextu...

Mixtral 8X22b Instruct V0.1

22B
by Mistral AI · 65,536 ctx

Mixtral 8X7b V0.1

7B
by Mistral AI · 32,768 ctx

Mixtral-8x7B Instruct v0.1

7B
by Mistral AI · 32,768 ctx

Molmo 7B D 0924

7B
by Allen Institute for AI (AI2) · 4,096 ctx
?

MoonshotAI: Kimi K2 0905

1000B
by Moonshot AI · 262,144 ctx

Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model d...

?

Morph: Morph V3 Fast

7B
by Morph · 81,920 ctx

Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the p...

?

Morph: Morph V3 Large

70B
by Morph · 262,144 ctx

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model...

?

MythoMax 13B

13B
by Gryphe · 4,096 ctx

One of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay. #merge

NVIDIA-Nemotron-3-Super-120B-A12B

120B
by Nvidia · 262,144 ctx

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent ...

NVIDIA: Llama 3.3 Nemotron Super 49B V1.5

49B
by Nvidia · 131,072 ctx

Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct wit...

NVIDIA: Nemotron 3 Nano 30B A3B

30B
by Nvidia · 262,144 ctx

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build special...

NVIDIA: Nemotron 3 Super

120B
by Nvidia · 1,000,000 ctx

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accu...

NVIDIA: Nemotron Nano 9B V2

9B
by Nvidia · 131,072 ctx

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reaso...

Nemotron-3-Nano-Omni-30B-A3B-Reasoning

30B
by Nvidia · 262,144 ctx

Nemotron 3 Nano Omni is an open multimodal model built on a hybrid Mixture-of-Experts (MoE) architecture, engineered for high efficiency ...

?

Nex AGI: DeepSeek V3.1 Nex N1

text
by Nex Agi · 131,072 ctx

DeepSeek V3.1 Nex-N1 is the flagship release of the Nex-N1 series — a post-trained model designed to highlight agent autonomy, tool use, ...

?

Nous Hermes 2 Mixtral 8X7B Dpo

7B
by Nous Research · 32,768 ctx
?

Nous: Hermes 3 405B Instruct

405B
by Nous Research · 131,072 ctx

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better rolepl...

?

Nous: Hermes 4 405B

405B
by Nous Research · 131,072 ctx

Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mo...

?

Nous: Hermes 4 70B

70B
by Nous Research · 131,072 ctx

Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the large...

?

NousResearch: Hermes 2 Pro - Llama-3 8B

8B
by Nous Research · 8,192 ctx

Hermes 2 Pro is an upgraded, retrained version of Nous Hermes 2, consisting of an updated and cleaned version of the OpenHermes 2.5 Datas...

Nvidia Nemotron Nano 9B V2

9B
by Nvidia · 131,072 ctx
?

Perceptron: Perceptron Mk1

multimodal
by Perceptron · 32,768 ctx

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and ...

?

Prime Intellect: INTELLECT-3

text
by Prime Intellect · 131,072 ctx

INTELLECT-3 is a 106B-parameter Mixture-of-Experts model (12B active) post-trained from GLM-4.5-Air-Base using supervised fine-tuning (SF...

Qwen 2 (1.5B)

5B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2 (72B)

72B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2 (7B)

7B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2 Instruct (1.5B)

5B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2.5 14B Instruct

14B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen 2.5 Coder 32B Instruct

32B
by Alibaba (Qwen Team) · 16,384 ctx

Qwen QwQ-32B

32B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2 72B Instruct

72B
by Togethercomputer · 32,768 ctx

Qwen2-VL (72B) Instruct

72B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 1.5B

5B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 1.5B Instruct

5B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 14B

14B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 32B

32B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 32B Instruct

32B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 3B Instruct

3B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 72B

72B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 72B Instruct

72B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 72B Instruct Turbo

72B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 7B

7B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 7B Instruct

7B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen2.5 7B Instruct Turbo

7B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 0.6B

6B
by Alibaba (Qwen Team) · 40,960 ctx

Qwen3 0.6B Base

6B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 1.7B

7B
by Alibaba (Qwen Team) · 40,960 ctx

Qwen3 1.7B Base

7B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 14B Base

14B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 30B A3b Base

30B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 4B Base

4B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3 4B Instruct 2507

4B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 8B Base

8B
by Alibaba (Qwen Team) · 32,768 ctx

Qwen3-235B-A22B-Instruct-2507

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-235B-A22B-Instruct-2507 is the updated version of the Qwen3-235B-A22B non-thinking mode, featuring Significant improvements in gene...

Qwen3.5-0.8B

8B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5-0.8B is Alibaba's smallest model in the Qwen3.5 series, featuring a hybrid Gated Delta Networks and sparse Mixture-of-Experts arc...

Qwen3.5-2B

2B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5-2B is a compact yet capable model from Alibaba's Qwen3.5 series. It features a 262K token context window, support for 201 languag...

Qwen3.5-4B

4B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5-4B is a mid-size model from Alibaba's Qwen3.5 series that delivers a strong balance of performance and efficiency. It features a ...

Qwen: Qwen Plus 0728

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, an...

Qwen: Qwen-Plus

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

Qwen: Qwen2.5 7B Instruct

7B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more...

Qwen: Qwen2.5 VL 72B Instruct

72B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing ...

Qwen: Qwen3 14B

14B
by Alibaba (Qwen Team) · 131,702 ctx

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialo...

Qwen: Qwen3 235B A22B

235B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supp...

Qwen: Qwen3 235B A22B Instruct 2507

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture...

Qwen: Qwen3 235B A22B Thinking 2507

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning ...

Qwen: Qwen3 30B A3B

30B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to e...

Qwen: Qwen3 30B A3B Instruct 2507

30B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. ...

Qwen: Qwen3 30B A3B Thinking 2507

30B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-st...

Qwen: Qwen3 32B

32B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dial...

Qwen: Qwen3 8B

8B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dia...

Qwen: Qwen3 Coder 30B A3B Instruct

30B
by Alibaba (Qwen Team) · 160,000 ctx

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed f...

Qwen: Qwen3 Coder 480B A35B

480B
by Alibaba (Qwen Team) · 1,048,576 ctx

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agenti...

Qwen: Qwen3 Coder Flash

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model...

Qwen: Qwen3 Coder Next

480B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse Mo...

Qwen: Qwen3 Coder Plus

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializ...

Qwen: Qwen3 Max

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual ...

Qwen: Qwen3 Max Thinking

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi...

Qwen: Qwen3 Next 80B A3B Instruct

80B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thi...

Qwen: Qwen3 Next 80B A3B Thinking

80B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. ...

Qwen: Qwen3 VL 235B A22B Instruct

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across image...

Qwen: Qwen3 VL 235B A22B Thinking

235B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. ...

Qwen: Qwen3 VL 30B A3B Instruct

30B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its ...

Qwen: Qwen3 VL 30B A3B Thinking

30B
by Alibaba (Qwen Team) · 131,072 ctx

Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its ...

Qwen: Qwen3 VL 32B Instruct

32B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across te...

Qwen: Qwen3 VL 8B Instruct

8B
by Alibaba (Qwen Team) · 256,000 ctx

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning ...

Qwen: Qwen3 VL 8B Thinking

8B
by Alibaba (Qwen Team) · 256,000 ctx

Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual rea...

Qwen: Qwen3.5 397B A17B

397B
by Alibaba (Qwen Team) · 262,144 ctx

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism ...

Qwen: Qwen3.5 Plus 2026-02-15

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with...

Qwen: Qwen3.5 Plus 2026-04-20

235B
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces t...

Qwen: Qwen3.5-122B-A10B

122B
by Alibaba (Qwen Team) · 262,144 ctx

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a ...

Qwen: Qwen3.5-27B

27B
by Alibaba (Qwen Team) · 262,144 ctx

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balanc...

Qwen: Qwen3.5-35B-A3B

35B
by Alibaba (Qwen Team) · 262,144 ctx

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechani...

Qwen: Qwen3.5-9B

9B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understandi...

Qwen: Qwen3.5-Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sp...

Qwen: Qwen3.6 27B

27B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid mult...

Qwen: Qwen3.6 35B A3B

35B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters pe...

Qwen: Qwen3.6 Flash

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M toke...

Qwen: Qwen3.6 Plus

multimodal
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling s...

Qwen: Qwen3.7 Max

text
by Alibaba (Qwen Team) · 1,000,000 ctx

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric worklo...

?

ReMM SLERP 13B

13B
by Undi95 · 6,144 ctx

A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge

Reka Edge

7B
by Rekaai · 16,384 ctx

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. ...

Reka Flash 3

21B
by Rekaai · 65,536 ctx

Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at gen...

?

Relace: Relace Apply 3

text
by Relace · 256,000 ctx

Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates fr...

?

Relace: Relace Search

text
by Relace · 256,000 ctx

The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user re...

?

Sao10K: Llama 3 8B Lunaris

8B
by Sao10k · 8,192 ctx

Lunaris 8B is a versatile generalist and roleplaying model based on Llama 3. It's a strategic merge of multiple models, designed to balan...

?

Sao10K: Llama 3.1 70B Hanami x1

70B
by Sao10k · 16,000 ctx

This is [Sao10K](/sao10k)'s experiment over [Euryale v2.2](/sao10k/l3.1-euryale-70b).

?

Sao10K: Llama 3.1 Euryale 70B v2.2

70B
by Sao10k · 131,072 ctx

Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3...

?

Sao10K: Llama 3.3 Euryale 70B

70B
by Sao10k · 131,072 ctx

Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B ...

?

Sao10k: Llama 3 Euryale 70B v2.1

70B
by Sao10k · 8,192 ctx

Euryale 70B v2.1 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). - Better prompt adherence. - Better ana...

Sarvam M

24B
by Sarvamai · 32,768 ctx
?

Seed-1.8

200B
by Bytedance · 256,000 ctx

Optimized specifically for multimodal agent scenarios. It features enhanced agent capabilities, upgraded multimodal comprehension, and mo...

?

Seed-2.0-code

multimodal
by Bytedance · 256,000 ctx

A coding model optimized for real-world development environments, with reliable tool use in common IDEs such as Claude Code. It delivers ...

?

Seed-2.0-pro

multimodal
by Bytedance · 256,000 ctx

Built for the Agent era, it delivers stable performance in complex reasoning and long-horizon tasks, including multi-step planning, visua...

StepFun: Step 3.5 Flash

199B
by Stepfun · 262,144 ctx

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it select...

?

Switchpoint Router

text
by Switchpoint · 131,072 ctx

Switchpoint AI's router instantly analyzes your request and directs it to the optimal AI from an ever-evolving library. As the world of L...

Tencent: Hunyuan A13B Instruct

80B
by Tencent · 131,072 ctx

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B ...

?

TheDrummer: Cydonia 24B V4.1

24B
by Thedrummer · 131,072 ctx

Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.

?

TheDrummer: Rocinante 12B

12B
by Thedrummer · 32,768 ctx

Rocinante 12B is designed for engaging storytelling and rich prose. Early testers have reported: - Expanded vocabulary with unique and ex...

?

TheDrummer: Skyfall 36B V2

36B
by Thedrummer · 32,768 ctx

Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-pla...

?

TheDrummer: UnslopNemo 12B

12B
by Thedrummer · 32,768 ctx

UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.

Tongyi DeepResearch 30B A3B

30B
by Alibaba (Qwen Team) · 131,072 ctx

Tongyi DeepResearch is an agentic large language model developed by Tongyi Lab, with 30 billion total parameters activating only 3 billio...

Upstage: Solar Pro 3

text
by Upstage · 128,000 ctx

Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forwa...

WizardLM-2 8x22B

22B
by Microsoft · 65,536 ctx

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprie...

Writer: Palmyra X5

70B
by Writer · 1,040,000 ctx

Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-...

Xiaomi: MiMo-V2-Flash

7B
by Xiaomi · 262,144 ctx

MiMo-V2-Flash is an open-source foundation language model developed by Xiaomi. It is a Mixture-of-Experts model with 309B total parameter...

Xiaomi: MiMo-V2-Omni

multimodal
by Xiaomi · 262,144 ctx

MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It comb...

Xiaomi: MiMo-V2-Pro

text
by Xiaomi · 1,048,576 ctx

MiMo-V2-Pro is Xiaomi's flagship foundation model, featuring over 1T total parameters and a 1M context length, deeply optimized for agent...

Xiaomi: MiMo-V2.5

315B
by Xiaomi · 1,048,576 ctx

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surp...

Xiaomi: MiMo-V2.5-Pro

65B
by Xiaomi · 1,048,576 ctx

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, an...

Z.ai: GLM 4 32B

32B
by Zhipu AI · 128,000 ctx

GLM 4 32B is a cost-effective foundation language model. It can efficiently perform complex tasks and has significantly enhanced capabili...

Z.ai: GLM 4.5V

9B
by Zhipu AI · 65,536 ctx

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 1...

Z.ai: GLM 4.6V

108B
by Zhipu AI · 131,072 ctx

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents,...

Z.ai: GLM 4.7 Flash

31B
by Zhipu AI · 202,752 ctx

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agenti...

Z.ai: GLM 5V Turbo

multimodal
by Zhipu AI · 202,752 ctx

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively ...

gemini-3.1-pro

multimodal
by Google DeepMind · 1,000,000 ctx

Bring any idea to life with state-of-the-art reasoning to help you learn, build, and plan anything. Best for complex tasks and bringing c...

?

inclusionAI: Ling-2.6-1T

text
by Inclusionai · 262,144 ctx

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents...

?

inclusionAI: Ling-2.6-flash

text
by Inclusionai · 262,144 ctx

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-w...

?

inclusionAI: Ring-2.6-1T

text
by Inclusionai · 262,144 ctx

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both str...

meta-llama/Llama-2-7b-chat-hf

7B
by Meta AI · 4,096 ctx

meta-llama/Llama-3.3-70B-Instruct

70B
by Meta AI · 131,072 ctx

nim/meta/llama-3.1-70b-instruct

70B
by Meta AI · 16,384 ctx

nim/meta/llama-3.1-8b-instruct

8B
by Meta AI · 16,384 ctx

nim/meta/llama-3.2-11b-vision-instruct

11B
by Nvidia · 16,384 ctx

nim/meta/llama-3.2-90b-vision-instruct

90B
by Meta AI · 16,384 ctx

nim/meta/llama-3.3-70b-instruct

70B
by Meta AI · 16,384 ctx

nim/mistralai/mixtral-8x22b-instruct-v01

22B
by Mistral AI · 16,384 ctx

nim/mistralai/mixtral-8x7b-instruct-v01

7B
by Mistral AI · 16,384 ctx

nim/nv-mistralai/mistral-nemo-12b-instruct

12B
by Nvidia · 16,384 ctx

nim/nvidia/llama-3.1-nemotron-70b-instruct

70B
by Nvidia · 16,384 ctx

nim/nvidia/llama-3.3-nemotron-super-49b-v1

49B
by Nvidia · 16,384 ctx
?

Kimi K2.6

1000B
by Moonshot AI · Kimi · 256,000 ctx

Long-horizon coding + autonomous-execution upgrade over K2.5.

?

Kimi K2.5

1000B
by Moonshot AI · Kimi · 256,000 ctx

Multimodal agentic variant — adds a vision encoder to the K2 backbone.

?

Kimi K2 Thinking

1000B
by Moonshot AI · Kimi · 256,000 ctx

Moonshot's open-weight reasoning variant — extended chain-of-thought training on top of Kimi K2.

GPT-OSS 120B

120B
by OpenAI · GPT-OSS · 128,000 ctx

OpenAI's first open-weight LLM in years — Apache-licensed MoE.

GPT-OSS 20B

20B
by OpenAI · GPT-OSS · 128,000 ctx

20B GPT-OSS — single-GPU local target.

GLM-4.5

355B
by Zhipu AI · GLM · 128,000 ctx

Zhipu's frontier open-weight MoE — 355B total, 32B active. Strong agentic + reasoning marks for an open model.

GLM-4.5-Air

106B
by Zhipu AI · GLM · 128,000 ctx

Smaller, cheaper sibling of GLM-4.5. 106B total, 12B active.

?

Kimi K2

1000B
by Moonshot AI · Kimi · 256,000 ctx

Moonshot's frontier open-weight MoE — 1T total, 32B active.

Qwen 3 14B

15B
by Alibaba (Qwen Team) · Qwen 3 · 128,000 ctx

Qwen 3 235B

235B
by Alibaba (Qwen Team) · Qwen 3 · 128,000 ctx

Alibaba's frontier MoE — 235B total / 22B active.

Qwen 3 32B

33B
by Alibaba (Qwen Team) · Qwen 3 · 128,000 ctx

Dense 32B Qwen 3.

Qwen 3 4B

4B
by Alibaba (Qwen Team) · Qwen 3 · 32,768 ctx

Qwen 3 8B

8B
by Alibaba (Qwen Team) · Qwen 3 · 128,000 ctx

Gemma 3 12B

12B
by Google DeepMind · Gemma · 128,000 ctx

12B Gemma 3 — multimodal, single-GPU target.

Gemma 3 1B

1B
by Google DeepMind · Gemma · 32,768 ctx

1B Gemma 3 — edge / mobile.

Gemma 3 27B

27B
by Google DeepMind · Gemma · 128,000 ctx

Google's open-weight multimodal LLM — efficient and license-permissive.

Gemma 3 4B

4B
by Google DeepMind · Gemma · 128,000 ctx

4B Gemma 3 — laptop multimodal.

DeepSeek R1

671B
by DeepSeek · DeepSeek · 128,000 ctx

DeepSeek's reasoning model — RL-trained, frontier-class, MIT-licensed.

DeepSeek R1 Distill Llama 70B

70B
by DeepSeek · DeepSeek · 128,000 ctx

70B Llama distilled from DeepSeek R1's reasoning traces.

DeepSeek R1 Distill Qwen 1.5B

2B
by DeepSeek · DeepSeek · 128,000 ctx

Tiny distilled R1 — phone / browser deployable.

DeepSeek R1 Distill Qwen 14B

15B
by DeepSeek · DeepSeek · 128,000 ctx

14B distilled R1 — laptop-friendly reasoning.

DeepSeek R1 Distill Qwen 32B

33B
by DeepSeek · DeepSeek · 128,000 ctx

32B Qwen base distilled from DeepSeek R1.

DeepSeek R1 Distill Qwen 7B

8B
by DeepSeek · DeepSeek · 128,000 ctx

7B distilled R1 — runs on any modern GPU.

MiniMax-Text-01

456B
by MiniMax · MiniMax-Text · 4,000,000 ctx

MiniMax's first open MoE — 456B total, 45.9B active. 1M+ context via Lightning Attention.

OLMo 3 7B

7B
by Allen Institute for AI (AI2) · OLMo · 4,096 ctx

Allen AI's latest fully-open OLMo — model + training data + checkpoints.

DeepSeek V3

671B
by DeepSeek · DeepSeek · 128,000 ctx

DeepSeek's flagship MoE — 671B total, 37B active, frontier-class.

IBM Granite 3.1 8B

8B
by IBM Research · Granite · 128,000 ctx

IBM's enterprise-focused open-weight LLM.

Phi-4

15B
by Microsoft · Phi · 16,384 ctx

Microsoft's 14B small-LM workhorse — punches above its weight.

Llama 3.3 70B

70B
by Meta AI · Llama · 128,000 ctx

Meta's best-in-class open-weight LLM — 70B class.

Qwen 2.5 Coder 32B

33B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

Alibaba's open-weight coding model — best in class for 32B.

Hunyuan-Large

389B
by Tencent · Hunyuan · 256,000 ctx

Tencent's open-weight MoE — 389B total, 52B active. Largest open MoE at launch.

Stable Diffusion 3.5 Medium

3B
by Stability AI · Stable Diffusion

2.5B SD 3.5 — fits on 12 GB consumer GPUs.

Stable Diffusion 3.5 Large

8B
by Stability AI · Stable Diffusion

Stability AI's latest open-weight image-gen — 8.1B params, MMDiT architecture.

Yi-Lightning

text
by 01.AI · Yi · 16,000 ctx

01.AI's fastest production model. Tops the LMSYS Arena Chinese leaderboard.

Llama 3.2 11B Vision

11B
by Meta AI · Llama · 128,000 ctx

Meta's open-weight multimodal LLM — vision + text in 11B.

Llama 3.2 1B

1B
by Meta AI · Llama · 128,000 ctx

Meta's smallest Llama — mobile + on-device target.

Llama 3.2 3B

3B
by Meta AI · Llama · 128,000 ctx

3B Llama — laptop-class chat + RAG.

Llama 3.2 90B Vision

90B
by Meta AI · Llama · 128,000 ctx

Meta's largest vision-capable Llama.

Qwen 2.5 14B

15B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

14B Qwen 2.5 — sweet spot for single-GPU local hosting.

Qwen 2.5 32B

33B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

32B Qwen 2.5 — laptop-class workhorse.

Qwen 2.5 3B

3B
by Alibaba (Qwen Team) · Qwen · 32,768 ctx

3B Qwen 2.5 — laptop / edge target.

Qwen 2.5 72B

73B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

Alibaba's flagship open-weight LLM — 72B dense.

Qwen 2.5 7B

8B
by Alibaba (Qwen Team) · Qwen · 128,000 ctx

7B Qwen 2.5 — most popular Qwen variant on Ollama.

Command R+

104B
by Cohere · Command · 128,000 ctx

Cohere's open-weight RAG-optimized LLM — multilingual + tool use.

Phi-3.5 Mini

4B
by Microsoft · Phi · 128,000 ctx

3.8B Phi — laptop / edge target.

?

Hermes 3 70B

70B
by Nous Research · Hermes · 128,000 ctx

Nous Research's flagship Llama fine-tune — agent-friendly.

?

Hermes 3 8B

8B
by Nous Research · Hermes · 128,000 ctx

FLUX.1 Dev

12B
by Black Forest Labs · FLUX

Open-weight FLUX.1 — non-commercial license.

FLUX.1 Pro

12B
by Black Forest Labs · FLUX

Black Forest Labs' flagship image-gen model — closed/API.

FLUX.1 Schnell

12B
by Black Forest Labs · FLUX

Distilled fast FLUX.1 — Apache-2.0, commercial-friendly.

Gemma 2 2B

3B
by Google DeepMind · Gemma 2 · 8,192 ctx

Tiny 2B Gemma 2 — laptop / mobile.

Mistral Large 2

123B
by Mistral AI · Mistral Large · 128,000 ctx

Mistral's flagship open-weight model — 123B dense.

Llama 3.1 405B

405B
by Meta AI · Llama · 128,000 ctx

Meta's largest open-weight LLM — dense 405B, frontier-class at launch.

Llama 3.1 70B

70B
by Meta AI · Llama · 128,000 ctx

Llama 3.1 70B — production workhorse, superseded by 3.3 but still widely deployed.

Llama 3.1 8B

8B
by Meta AI · Llama · 128,000 ctx

Meta's most popular open-weight small LLM — fits anywhere.

Mistral Nemo 12B

12B
by Mistral AI · Mistral · 128,000 ctx

Mistral × Nvidia collab — 12B Apache-licensed, multilingual.

InternLM 2.5 20B

20B
by Shanghai AI Lab · InternLM · 1,000,000 ctx

Shanghai AI Lab's dense 20B open-weight. Strong long-context + tool use for its size.

Gemma 2 27B

27B
by Google DeepMind · Gemma 2 · 8,192 ctx

Google's pre-Gemma-3 open-weight workhorse.

Gemma 2 9B

9B
by Google DeepMind · Gemma 2 · 8,192 ctx

9B Gemma 2 — single-GPU local target.

DeepSeek Coder V2 236B

236B
by DeepSeek · DeepSeek Coder · 128,000 ctx

DeepSeek's MoE coding model — 236B total, 21B active.

DeepSeek Coder V2 Lite

16B
by DeepSeek · DeepSeek Coder · 128,000 ctx

16B MoE / 2.4B active — laptop-class coder.

Mistral 7B v0.3

7B
by Mistral AI · Mistral 7B · 32,768 ctx

The current Mistral 7B — adds function calling + extended vocab.

IBM Granite Code 8B

8B
by IBM Research · Granite · 4,096 ctx

Phi-3 Medium

14B
by Microsoft · Phi · 128,000 ctx

Phi-3 Mini

4B
by Microsoft · Phi · 128,000 ctx

Mixtral 8x22B

141B
by Mistral AI · Mixtral · 65,536 ctx

Mistral's open MoE — 141B total, 39B active.

Mistral 7B v0.2

7B
by Mistral AI · Mistral 7B · 32,768 ctx

Mistral 7B v0.2 — earlier 32K context revision.

mxbai-embed-large

0B
by Mixedbread AI · Mixedbread Embed · 512 ctx

335M embedding model — top MTEB scores for its size.

Moondream 1.8B

2B
by Moondream · Moondream · 2,048 ctx

Tiny multimodal — laptop-class image understanding.

Nomic Embed Text

0B
by Nomic AI · Nomic Embed · 8,192 ctx

Open embedding model — 69M Ollama pulls, the local default.

OLMo 7B

7B
by Allen Institute for AI (AI2) · OLMo · 2,048 ctx

LLaVA 13B

13B
by LLaVA Project · LLaVA · 4,096 ctx

LLaVA 34B

34B
by LLaVA Project · LLaVA · 4,096 ctx

Largest open-weight LLaVA — vision encoder + Yi-34B backbone.

LLaVA 7B

7B
by LLaVA Project · LLaVA · 4,096 ctx
?

BGE-M3

1B
by BAAI (Beijing Academy of AI) · BGE · 8,192 ctx

Multilingual + multifunctional embedding (100+ languages).

Code Llama 70B

70B
by Meta AI · Code Llama · 16,384 ctx

Meta's largest code-specialised Llama.

TinyLlama 1.1B

1B
by TinyLlama Project · TinyLlama · 2,048 ctx

1.1B Llama-arch model — 3T training tokens.

Whisper Large v3

2B
by OpenAI · Whisper · 30 ctx

OpenAI's open-weight speech-to-text — the standard transcription model.

DeepSeek Coder 33B

33B
by DeepSeek · DeepSeek Coder · 16,384 ctx

DeepSeek Coder 6.7B

7B
by DeepSeek · DeepSeek Coder · 16,384 ctx

Yi-34B

34B
by 01.AI · Yi · 32,000 ctx

Earlier open-weight Yi release — bilingual EN/ZH, 32K-token context.

Mistral 7B v0.1

7B
by Mistral AI · Mistral 7B · 8,192 ctx

Original Mistral 7B — historical reference.

Baichuan2-13B

13B
by Baichuan Inc. · Baichuan · 4,096 ctx

Bilingual EN/ZH open-weight. Strong for its size on Chinese-language benchmarks.

Code Llama 13B

13B
by Meta AI · Code Llama · 16,384 ctx

Code Llama 34B

34B
by Meta AI · Code Llama · 16,384 ctx

Code Llama 7B

7B
by Meta AI · Code Llama · 16,384 ctx

Stable Diffusion XL

4B
by Stability AI · Stable Diffusion

Workhorse open-weight image-gen — 3.5B params, runs anywhere.

Whisper Base

0B
by OpenAI · Whisper · 30 ctx

74M Whisper — browser / Raspberry Pi-deployable.

Whisper Medium

1B
by OpenAI · Whisper · 30 ctx

769M Whisper variant — half the size of Large, 80% of the accuracy.

Whisper Small

0B
by OpenAI · Whisper · 30 ctx

244M Whisper — fits on edge GPUs and CPU.

Whisper Tiny

0B
by OpenAI · Whisper · 30 ctx

39M Whisper — runs in-browser via WebGPU.

Stable Diffusion 1.5

1B
by Stability AI · Stable Diffusion

The original viral image-gen model — still searched heavily.

Meta: Muse Spark 1.3 Contributor

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and...

GLM 5.3 FP8

text
by Zhipu AI · 1,048,576 ctx

GLM 5.3 FP8 Lora

text
by Zhipu AI · 1,048,576 ctx
?

Inception: Mercury 2.5 Preview

text
by Inception · 260,000 ctx

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, ...

Tencent: Hy4 preview

text
by Tencent · 1,048,576 ctx

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding ...

DeepSeek: DeepSeek V4 Flash Vision Exp

multimodal
by DeepSeek · 1,048,576 ctx

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepsee...

Meta: Muse Spark 1.2 Contributor

multimodal
by Meta AI · 1,048,576 ctx

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. ...

GLM 5.2 FP8 Lora

text
by Zhipu AI · 1,048,576 ctx

GLM 5.2 FP8

text
by Zhipu AI · 1,048,576 ctx

gemma-4-31B-it-Ultra

31B
by Google DeepMind · 131,072 ctx

Ultra speed version of gemma-4-31B-it

gpt-oss-120b-Ultra

120B
by OpenAI · 131,072 ctx

Ultra fast version of gpt-oss-120b

?

Inkling FP4

952B
by Thinking Machines · 524,288 ctx

Qwen3.5 0.8B Lora

8B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 122B A10B Lora

122B
by Alibaba (Qwen Team) · 8,192 ctx

Qwen3.5 27B Lora

27B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 2B Lora

2B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 35B A3B Base Lora

35B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 397B A17B Lora

397B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 4B Lora

4B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 9B Lora

9B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6 27B Lora

27B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6 35B A3B Lora

35B
by Alibaba (Qwen Team) · 262,144 ctx
?

MiniMax-M2.7-Turbo

text
by Minimaxai · 196,608 ctx

Speed-optimized MiniMax-M2.7

NVIDIA Nemotron 3 Ultra NVFP4

text
by Nvidia · 262,144 ctx

Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, re...

NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16

550B
by Nvidia · 262,144 ctx

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It deli...

Qwen3.5 35B A3b LoRa

35B
by Alibaba (Qwen Team) · 262,144 ctx

GLM 5 Fp4

text
by Zhipu AI · 202,752 ctx

Glm 4.7 Fp8

text
by Zhipu AI · 202,752 ctx

Llama 4 Maverick 17B 128E Instruct Nvfp4

17B
by Meta AI · 1,048,576 ctx

GLM 4.7 FP4

text
by Togethercomputer · 202,752 ctx

Kimi K2.5 FP4

text
by Togethercomputer · 262,144 ctx

MiniMax M2.5 FP4

text
by MiniMax · 8,192 ctx

Cogito V1 Preview Llama 70B

70B
by Deepcogito · 131,072 ctx

Cogito V1 Preview Llama 70B Turbo

70B
by Deepcogito · 131,072 ctx

Cogito V1 Preview Llama 8B

8B
by Deepcogito · 131,072 ctx

Cogito V1 Preview Qwen 14B

14B
by Deepcogito · 131,072 ctx

Cogito V1 Preview Qwen 32B

32B
by Deepcogito · 131,072 ctx

DeepSeek: DeepSeek V3.2 Exp

671B
by DeepSeek · 163,840 ctx

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectu...

Deepcoder 14B Preview

14B
by Togethercomputer · 131,072 ctx

Gemma 3 270M It Lora

0B
by Google DeepMind · 32,768 ctx

Gemma 3 27B It Lora

27B
by Google DeepMind

Gemma 4 31B It Lora

31B
by Google DeepMind · 262,144 ctx

Glm 4.5 Air Fp8

text
by Zhipu AI · 131,072 ctx
?

L3-8B-Lunaris-v1-Turbo

8B
by Sao10k · 8,192 ctx

Llama 3.3 70B Instruct FP8 Lora

70B
by Meta AI · 131,072 ctx

Llama 4 Maverick Instruct (17Bx128E) FP8

17B
by Meta AI · 1,048,576 ctx

Llama 4 Scout 17B 16E Instruct Fp8 Lora

17B
by Meta AI · 10,485,760 ctx

Meta Llama 3.1 8B Instruct Awq Int4

8B
by Meta AI · 131,072 ctx

Mixtral 8x7B Instruct V0.1 FP8 Lora

7B
by Mistral AI · 32,768 ctx
?

MoonshotAI Kimi Latest

1000B
by Moonshot AI · 262,144 ctx

This model always redirects to the latest model in the MoonshotAI Kimi family.

Nemotron 3 Nano Omni 30B A3b Reasoning Fp8

30B
by Nvidia · 131,072 ctx

Nvidia Nemotron 3 Nano 30B A3b Bf16

30B
by Nvidia · 262,144 ctx

Nvidia Nemotron 3 Super 120B A12b Bf16

120B
by Nvidia · 262,144 ctx

Nvidia Nemotron 3 Super 120B A12b Fp8

120B
by Nvidia · 262,144 ctx

Qwen3 235B A22B Instruct 2507 FP8 Throughput

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 30B A3B Instruct 2507 Lora

30B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 8B Lora

8B
by Alibaba (Qwen Team) · 40,960 ctx

Qwen3 Coder 480B A35B Instruct Fp8

480B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 Coder Next Fp8

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3 Next 80B A3b Instruct Fp8

80B
by Alibaba (Qwen Team)

Qwen3-Coder-480B-A35B-Instruct-Turbo

480B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3-Coder-480B-A35B-Instruct is the Qwen3's most agentic code model, featuring Significant Performance on Agentic Coding, Agentic Brows...

Qwen3-VL-235B-A22B-Instruct-FP8

235B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 122B A10b Fp8

122B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.5 9B Fp8

9B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6 35B A3b Fp8

35B
by Alibaba (Qwen Team) · 262,144 ctx

Qwen: Qwen3.6 Max Preview

text
by Alibaba (Qwen Team) · 262,144 ctx

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximate...

Tencent: Hy3 preview

299B
by Tencent · 262,144 ctx

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports con...

gemma-4-31B-it-turbo

31B
by Google DeepMind · 262,144 ctx

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating te...

gpt-oss-120b-Turbo

120B
by OpenAI · 131,072 ctx