Mercury 2.5 is a diffusion large language model that InceptionLabs launched on 8 September 2026 as the fastest reasoning model in production, generating 1,107 tokens per second. InceptionLabs prices Mercury 2.5 at $0.04 per million input tokens and $0.15 per million output tokens, and rates its quality as comparable to Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna.

InceptionLabs (brand name: Inception) develops the Mercury family of diffusion language models. The company announced Mercury 2.5 through its official blog on 8 September 2026, alongside two preview products: Mercury Voice, a speech model with a time-to-first-token under 170 milliseconds, and Mercury Router, a model-routing service. Mercury 2.5 is available through the Inception API, OpenRouter, and Baseten. The launch matters to business buyers because the fastest AI model of 2026 now costs a fraction of the transformer models that most AI tools run on.

  • Speed: 1,107 tokens per second in production, according to InceptionLabs
  • Intelligence: 40% higher than Mercury 2, the previous release, according to InceptionLabs
  • Price: $0.04 per million input tokens, $0.15 per million output tokens
  • Context window: 260,000 tokens
  • Quality tier: comparable to Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna (Low), per Inception’s own benchmarks

Mercury 2.5 Specifications: Speed, Price, and Context

Mercury 2.5 delivers 1,107 tokens per second, a 260,000-token context window, and a price of $0.04 per million input tokens and $0.15 per million output tokens, according to the official InceptionLabs announcement and the model listing on BenchLM. The table below lists the 5 published attributes of Mercury 2.5.

Attribute Mercury 2.5 Source
Output speed 1,107 tokens/second InceptionLabs blog
Intelligence gain vs Mercury 2 +40% InceptionLabs blog
Context window 260,000 tokens BenchLM
Input price $0.04 per 1M tokens BenchLM
Output price $0.15 per 1M tokens BenchLM

InceptionLabs positions Mercury 2.5 against 3 cost-optimized frontier models: Claude Haiku 4.5 from Anthropic, Gemini 3.5 Flash-Lite from Google, and GPT-5.6 Luna (Low) from OpenAI, according to the launch announcement distributed via Yahoo Finance. Developers can access Mercury 2.5 today through the OpenRouter listing as a preview model.

How a Diffusion LLM Differs From GPT, Claude, and Gemini

A diffusion LLM is a language model that generates text through iterative denoising, refining a whole block of tokens in parallel over repeated passes, whereas transformer models such as GPT-5.6, Claude Haiku 4.5, and Gemini 3.5 generate one token after another. Parallel generation is the source of the speed difference: Mercury 2.5 outputs 1,107 tokens per second, while comparable-quality transformer models typically output 60 to 120 tokens per second.

What Mercury 2.5 Changes for Businesses Choosing AI Tools

Mercury 2.5 sets a new cost floor for the vendors that build AI writing assistants, customer-service chatbots, and marketing-automation tools, because a vendor that switches its underlying model to Mercury 2.5 pays $0.04 per million input tokens instead of the higher rates of transformer models at the same quality tier. Businesses evaluating AI tools now have 3 concrete questions to ask a vendor: which model powers the product, how fast responses arrive, and whether falling model costs reach the per-seat price.

Speed changes the product experience as well as the bill: a writing tool running on Mercury 2.5 drafts a 1,000-word document in about 1 second. Buyers comparing vendors on our list of the 15 Best AI Tools for Business in 2026 should treat response latency and model cost as evaluation criteria alongside features, because the Mercury 2.5 launch makes both measurable and negotiable.

What Is Not Yet Verified About Mercury 2.5

The claim that Mercury 2.5 matches Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna in quality is InceptionLabs’ own benchmark result, and no independent third party had published a verification as of 9 September 2026. The Hacker News discussion of the launch, which reached 212 points, centered on 2 open questions: whether diffusion LLMs match transformer quality on real tasks, and whether Inception’s benchmark selection reflects general performance. Mercury Voice and Mercury Router are preview products, not generally available services. InceptionLabs has not disclosed its team size or funding in the launch materials.

💡 Our Take: The 2026 AI race has split into two contests: raw intelligence at the top, and “good enough” intelligence at the lowest latency and cost. Mercury 2.5 is the sharpest move yet in the second contest. For businesses buying AI tools, the practical consequence is simple: vendor pricing should be falling this year, and a vendor whose price is not falling owes you an explanation.

Frequently Asked Questions

What is Mercury 2.5?

Mercury 2.5 is a diffusion large language model from InceptionLabs, released on 8 September 2026. InceptionLabs reports it generates 1,107 tokens per second, offers a 260,000-token context window, and delivers 40% higher intelligence than Mercury 2.

How much does Mercury 2.5 cost?

Mercury 2.5 costs $0.04 per million input tokens and $0.15 per million output tokens, according to the BenchLM model listing. The model is available through the Inception API, OpenRouter, and Baseten.

Is Mercury 2.5 the fastest AI model in 2026?

Yes, according to InceptionLabs, Mercury 2.5 is the fastest reasoning large language model in production as of September 2026, at 1,107 tokens per second. Comparable-quality transformer models typically generate 60 to 120 tokens per second.

Does Mercury 2.5 beat Claude Haiku 4.5?

No published source claims Mercury 2.5 beats Claude Haiku 4.5. InceptionLabs describes Mercury 2.5 quality as comparable to Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna (Low), and that comparison rests on Inception’s own benchmarks without independent verification as of 9 September 2026.

What is a diffusion LLM?

A diffusion LLM is a language model that produces text by iteratively denoising a block of tokens in parallel, instead of generating one token at a time as transformer models do. The parallel process is what allows Mercury 2.5 to reach 1,107 tokens per second.

For Context: The Race to Cheaper, Faster AI Models

Cheaper inference is the trend across every model tier in 2026, from Mercury 2.5 at the low end to Claude Opus 5 at half the price of Fable 5 at the frontier. Buyers weighing an agent deployment should also read our coverage of DeepSeek V4-Flash on agent benchmarks, the other low-cost model that vendors cite in 2026 pricing negotiations.

Share.

I am a software engineer, I have a passion for working with cutting-edge technologies and staying up-to-date with the latest developments in the field. In my articles, I share my knowledge and insights on a range of topics, including business software, how to set up tools, and the latest trends in the tech industry.

Comments are closed.

Exit mobile version