by Inception

Inception: Mercury 2.5.

text open weights 260K ctx
Cheapest input
$0.04/M
on OpenRouter
Cheapest output
$0.15/M
on OpenRouter
Hosted equiv.
~$0.05/hr
@ 100 tok/s on OpenRouter

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

Where to use it

Cheapest hosted endpoints.

Provider Access $/M in $/M out
OpenRouter api aggregator $0.04 $0.15 Launch ↗
FAQ

Frequently asked.

How do I run Inception: Mercury 2.5?
Inception: Mercury 2.5 is open-weight, so you can self-host on rented GPUs. See the Run It Yourself tab for GPU configurations + cost estimates, or use one of the hosted inference providers listed on this page.
Where can I access Inception: Mercury 2.5?
Inception: Mercury 2.5 is available via OpenRouter. Each access option lists its own pricing (per million tokens or hourly hosting).
How much does it cost to run Inception: Mercury 2.5?
API pricing starts at $0.04/M input tokens and $0.15/M output tokens. Self-hosting cost depends on the GPU you rent — see the Run It Yourself tab.
Is Inception: Mercury 2.5 open-source or proprietary?
Inception: Mercury 2.5 is open-weight under the license. You can download and self-host it.