Published:

DeepSeek released V4 Pro 0813 on August 12, 2026 — its 1.6 trillion-parameter mixture-of-experts flagship — as the general-availability exit from a four-month preview period, with no official announcement, no blog post, and no product page, discoverable only through an OpenRouter listing and benchmark figures circulated in the DeepSeek WeChat group. Agentic coding benchmarks jumped sharply from the preview version: DeepSWE from 12.8 to 62.7, CyberGym from 52.7 to 83.3, and Terminal Bench 2.1 from 72.1 to 87.9 — placing V4 Pro 0813 near Fable 5 territory at $0.87/M output tokens, roughly 1/60th the cost of Fable 5, according to analysis by DigitalApplied. DeepSeek has also warned that a “significant” price increase is coming, with no date or new rate published.

What Changed: V4 Pro 0813 vs. the April Preview

DeepSeek V4 Pro 0813 is not a new product — it is the GA date-stamped release of the same MoE architecture first published in preview four months ago, with updated weights producing substantially different benchmark results. DeepSWE, which measures agentic software-engineering capability, rose from 12.8 to 62.7. CyberGym, which tests agentic security-task performance, rose from 52.7 to 83.3. Terminal Bench 2.1, an agentic coding benchmark run in a headless terminal environment, rose from 72.1 to 87.9. These figures originated from the DeepSeek WeChat group and were subsequently reported by Simon Willison at simonwillison.net and circulated on Reddit’s r/LocalLLaMA; they are preliminary, not publications from DeepSeek directly.

DeepSeek V4 Pro is a 1.6 trillion-total-parameter model with approximately 49 billion active parameters per forward pass (mixture-of-experts architecture), a 1 million-token context window, and up to 384K output tokens — confirmed via the OpenRouter listing at openrouter.ai/deepseek/deepseek-v4-pro-0813.

V4 Pro 0813 is a different model from DeepSeek V4-Flash, the smaller, faster sibling optimized for cost-per-task on bounded workloads. V4 Pro is the full MoE flagship with the larger context window and the higher benchmark scores.

Pricing: $0.87/M Output — with a Significant Hike on the Way

API pricing for DeepSeek V4 Pro 0813 is $0.435/M input tokens (cache miss), $0.003625/M input tokens (cache hit), and $0.87/M output tokens. The cache-hit rate makes repeat-context workloads — coding agents that re-read the same codebase headers on every call — exceptionally cheap: 10 million cached input tokens per month costs $36.25 at that rate.

DeepSeek has warned that a price increase is imminent. “DeepSeek says it plans to raise API prices in the near future and expects the increase to be significant, though it has not published the new rates or an effective date,” WCCFTech reported on August 12, 2026. The company has historically priced its API below cost as a market-penetration strategy; the acknowledgment of a coming hike signals the end of that subsidized phase.

Open Weights: April Preview Confirmed — 0813 Weights Unconfirmed

The April preview version of DeepSeek V4 Pro is available on HuggingFace at deepseek-ai/DeepSeek-V4-Pro. As of August 13, 2026, the 0813 GA weights have not been confirmed as published; API access via OpenRouter is the only confirmed path to the updated model. “I had to link to OpenRouter because DeepSeek don’t have any obvious announcement page for their new model,” wrote Simon Willison on simonwillison.net.

What V4 Pro 0813 Means for Business AI Buyers

For teams evaluating AI agents for business tasks, V4 Pro 0813 resets the price-performance floor for agentic coding. Benchmark parity with Fable 5 at approximately 1/60th the cost — if these preliminary DeepSWE, CyberGym, and Terminal Bench scores hold under production conditions — represents a material cost argument for any team currently paying frontier prices for coding-agent access. The pending price hike adds a concrete time-pressure element: DeepSeek has explicitly signaled that current rates are not permanent.

For businesses evaluating AI tools for business workflows that require large-context reasoning or multi-step agentic tasks, V4 Pro 0813’s 1 million-token context window and cache-hit economics are worth benchmarking directly against existing Fable 5 or GPT-5.6 Luna spend before the increase lands. Teams with high repeat-context usage — legal document analysis, large-codebase review, extended customer-service threads — stand to gain the most from the cache-hit pricing.

DeepSeek’s chip ambitions further context the pricing strategy: the company is building its own AI chip to reduce dependence on Nvidia, which would reduce its inference cost floor over time, even as it moves away from subsidized pricing on the current generation.

Our Take

DeepSeek continues to move the market by saying nothing — a quiet GA with near-Fable-5 agentic coding performance, announced via WeChat and reported by community observers rather than a press team. The coming price hike is the real signal for business buyers: it marks the end of DeepSeek’s subsidized phase, and teams that want below-cost frontier-class performance have a finite window to lock in current rates.


For Context — earlier witho2.com coverage of DeepSeek and related AI agents: DeepSeek V4-Flash beats its own flagship on agent benchmarks · DeepSeek’s chip ambitions. Related: open-weight AI agents entering the market alongside proprietary alternatives.

Share.

I am a software engineer, I have a passion for working with cutting-edge technologies and staying up-to-date with the latest developments in the field. In my articles, I share my knowledge and insights on a range of topics, including business software, how to set up tools, and the latest trends in the tech industry.

Comments are closed.

Exit mobile version