Close Menu
WithO2WithO2

    Subscribe to Updates

    Get the latest AI News Tools Updates in your Inbox

    What's Hot

    Shopify Acquires Tailwind Labs — New Paid Sign-Ups Closed as Revenue Fell 80%

    September 11, 2026

    Workflow Automation Software: How to Choose the Right Tool

    September 11, 2026

    Mercury 2.5 Hits 1,107 Tokens/Sec — Claude Haiku-Class AI at $0.04 per Million Tokens

    September 11, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    WithO2WithO2
    • AI
    • Blog
    • Business Software
    • Trending News
    • Stories
    WithO2WithO2
    Home » Trending News
    Trending News

    Mercury 2.5 Hits 1,107 Tokens/Sec — Claude Haiku-Class AI at $0.04 per Million Tokens

    By Amitabh SarkarSeptember 11, 20266 Mins Read0
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Mercury 2.5 diffusion language model generating tokens at high speed
    Mercury 2.5 generates 1,107 tokens per second through parallel diffusion, not one token at a time.
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Mercury 2.5 is a diffusion large language model that InceptionLabs launched on 8 September 2026 as the fastest reasoning model in production, generating 1,107 tokens per second. InceptionLabs prices Mercury 2.5 at $0.04 per million input tokens and $0.15 per million output tokens, and rates its quality as comparable to Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna.

    InceptionLabs (brand name: Inception) develops the Mercury family of diffusion language models. The company announced Mercury 2.5 through its official blog on 8 September 2026, alongside two preview products: Mercury Voice, a speech model with a time-to-first-token under 170 milliseconds, and Mercury Router, a model-routing service. Mercury 2.5 is available through the Inception API, OpenRouter, and Baseten. The launch matters to business buyers because the fastest AI model of 2026 now costs a fraction of the transformer models that most AI tools run on.

    • Speed: 1,107 tokens per second in production, according to InceptionLabs
    • Intelligence: 40% higher than Mercury 2, the previous release, according to InceptionLabs
    • Price: $0.04 per million input tokens, $0.15 per million output tokens
    • Context window: 260,000 tokens
    • Quality tier: comparable to Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna (Low), per Inception’s own benchmarks

    Table of Contents

    Toggle
    • Mercury 2.5 Specifications: Speed, Price, and Context
    • How a Diffusion LLM Differs From GPT, Claude, and Gemini
    • What Mercury 2.5 Changes for Businesses Choosing AI Tools
    • What Is Not Yet Verified About Mercury 2.5
    • Frequently Asked Questions
    • For Context: The Race to Cheaper, Faster AI Models

    Mercury 2.5 Specifications: Speed, Price, and Context

    Mercury 2.5 delivers 1,107 tokens per second, a 260,000-token context window, and a price of $0.04 per million input tokens and $0.15 per million output tokens, according to the official InceptionLabs announcement and the model listing on BenchLM. The table below lists the 5 published attributes of Mercury 2.5.

    AttributeMercury 2.5Source
    Output speed1,107 tokens/secondInceptionLabs blog
    Intelligence gain vs Mercury 2+40%InceptionLabs blog
    Context window260,000 tokensBenchLM
    Input price$0.04 per 1M tokensBenchLM
    Output price$0.15 per 1M tokensBenchLM

    InceptionLabs positions Mercury 2.5 against 3 cost-optimized frontier models: Claude Haiku 4.5 from Anthropic, Gemini 3.5 Flash-Lite from Google, and GPT-5.6 Luna (Low) from OpenAI, according to the launch announcement distributed via Yahoo Finance. Developers can access Mercury 2.5 today through the OpenRouter listing as a preview model.

    How a Diffusion LLM Differs From GPT, Claude, and Gemini

    A diffusion LLM is a language model that generates text through iterative denoising, refining a whole block of tokens in parallel over repeated passes, whereas transformer models such as GPT-5.6, Claude Haiku 4.5, and Gemini 3.5 generate one token after another. Parallel generation is the source of the speed difference: Mercury 2.5 outputs 1,107 tokens per second, while comparable-quality transformer models typically output 60 to 120 tokens per second.

    What Mercury 2.5 Changes for Businesses Choosing AI Tools

    Mercury 2.5 sets a new cost floor for the vendors that build AI writing assistants, customer-service chatbots, and marketing-automation tools, because a vendor that switches its underlying model to Mercury 2.5 pays $0.04 per million input tokens instead of the higher rates of transformer models at the same quality tier. Businesses evaluating AI tools now have 3 concrete questions to ask a vendor: which model powers the product, how fast responses arrive, and whether falling model costs reach the per-seat price.

    Speed changes the product experience as well as the bill: a writing tool running on Mercury 2.5 drafts a 1,000-word document in about 1 second. Buyers comparing vendors on our list of the 15 Best AI Tools for Business in 2026 should treat response latency and model cost as evaluation criteria alongside features, because the Mercury 2.5 launch makes both measurable and negotiable.

    What Is Not Yet Verified About Mercury 2.5

    The claim that Mercury 2.5 matches Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna in quality is InceptionLabs’ own benchmark result, and no independent third party had published a verification as of 9 September 2026. The Hacker News discussion of the launch, which reached 212 points, centered on 2 open questions: whether diffusion LLMs match transformer quality on real tasks, and whether Inception’s benchmark selection reflects general performance. Mercury Voice and Mercury Router are preview products, not generally available services. InceptionLabs has not disclosed its team size or funding in the launch materials.

    💡 Our Take: The 2026 AI race has split into two contests: raw intelligence at the top, and “good enough” intelligence at the lowest latency and cost. Mercury 2.5 is the sharpest move yet in the second contest. For businesses buying AI tools, the practical consequence is simple: vendor pricing should be falling this year, and a vendor whose price is not falling owes you an explanation.

    Frequently Asked Questions

    What is Mercury 2.5?

    Mercury 2.5 is a diffusion large language model from InceptionLabs, released on 8 September 2026. InceptionLabs reports it generates 1,107 tokens per second, offers a 260,000-token context window, and delivers 40% higher intelligence than Mercury 2.

    How much does Mercury 2.5 cost?

    Mercury 2.5 costs $0.04 per million input tokens and $0.15 per million output tokens, according to the BenchLM model listing. The model is available through the Inception API, OpenRouter, and Baseten.

    Is Mercury 2.5 the fastest AI model in 2026?

    Yes, according to InceptionLabs, Mercury 2.5 is the fastest reasoning large language model in production as of September 2026, at 1,107 tokens per second. Comparable-quality transformer models typically generate 60 to 120 tokens per second.

    Does Mercury 2.5 beat Claude Haiku 4.5?

    No published source claims Mercury 2.5 beats Claude Haiku 4.5. InceptionLabs describes Mercury 2.5 quality as comparable to Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna (Low), and that comparison rests on Inception’s own benchmarks without independent verification as of 9 September 2026.

    What is a diffusion LLM?

    A diffusion LLM is a language model that produces text by iteratively denoising a block of tokens in parallel, instead of generating one token at a time as transformer models do. The parallel process is what allows Mercury 2.5 to reach 1,107 tokens per second.

    For Context: The Race to Cheaper, Faster AI Models

    • OpenAI Ultrafast: GPT-5.6 Sol Now Runs 14× Faster via Cerebras — OpenAI’s transformer speed push reached 750 tokens per second on Cerebras hardware in August 2026
    • Gemini 3.8 Flash: Google’s Best Agent Model at $0.75/1M Tokens — Google’s September 2026 entry in the low-cost agent model tier
    • Claude Fable 5.1: 75% Cheaper Cache Cuts Agentic API Costs by 45% — Anthropic’s September 2026 price move on the frontier tier
    • Stripe Buys OpenRouter for $7B: What It Means for Your AI Stack — the model marketplace where Mercury 2.5 is listed changed hands in August 2026

    Cheaper inference is the trend across every model tier in 2026, from Mercury 2.5 at the low end to Claude Opus 5 at half the price of Fable 5 at the frontier. Buyers weighing an agent deployment should also read our coverage of DeepSeek V4-Flash on agent benchmarks, the other low-cost model that vendors cite in 2026 pricing negotiations.

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Amitabh Sarkar
    • Website

    I am a software engineer, I have a passion for working with cutting-edge technologies and staying up-to-date with the latest developments in the field. In my articles, I share my knowledge and insights on a range of topics, including business software, how to set up tools, and the latest trends in the tech industry.

    Related Posts

    Shopify Acquires Tailwind Labs — New Paid Sign-Ups Closed as Revenue Fell 80%

    September 11, 2026

    Meta Muse App Launches: Personal AI Agent at $20–$100/Month

    September 11, 2026

    OpenAI’s AI Agents Now Do a Researcher’s Job — What It Means

    September 10, 2026

    Comments are closed.

    Don't Miss
    Trending News

    Shopify Acquires Tailwind Labs — New Paid Sign-Ups Closed as Revenue Fell 80%

    By Amitabh SarkarSeptember 11, 2026

    Shopify acquired Tailwind Labs, the Canadian company behind the Tailwind CSS framework, on 9 September…

    Workflow Automation Software: How to Choose the Right Tool

    September 11, 2026

    Meta Muse App Launches: Personal AI Agent at $20–$100/Month

    September 11, 2026

    How Do AI Agents Work? The 5-Step Loop Explained

    September 11, 2026

    Subscribe to Updates

    Get the latest creative news from SmartMag about art & design.

    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    Our Picks

    Workflow Automation Software: How to Choose the Right Tool

    September 11, 2026

    Shopify vs WooCommerce vs BigCommerce 2026: Which Platform Wins?

    August 31, 2026

    12 Best Project Management Software Tools in 2026

    August 1, 2026

    9 Best Ecommerce Platforms Compared (2026)

    July 30, 2026
    Editors Picks

    Shopify Acquires Tailwind Labs — New Paid Sign-Ups Closed as Revenue Fell 80%

    September 11, 2026

    Meta Muse App Launches: Personal AI Agent at $20–$100/Month

    September 11, 2026

    OpenAI’s AI Agents Now Do a Researcher’s Job — What It Means

    September 10, 2026

    OpenAI Chief Scientist: AI Is Racing Toward Self-Improvement

    September 10, 2026
    About Us
    About Us

    Your Source for Innovation: Discover in-depth guides, solutions, and tools tailored to modern business challenges.

    Links
    • Blog
    • Privacy Policy
    • Contact WithO2.com
    • Terms and Conditions
    Facebook X (Twitter) Instagram Pinterest
    • About
    • Editorial Policy
    • Contact
    • Privacy Policy
    • Terms
    © 2026 WITHO2.COM

    Type above and press Enter to search. Press Esc to cancel.