Published: · Updated:
DeepSeek released V4 Pro 0813 on August 12, 2026 — its 1.6 trillion-parameter mixture-of-experts flagship — as the general-availability exit from a four-month preview period, with no official announcement, no blog post, and no product page, discoverable only through an OpenRouter listing and benchmark figures circulated in the DeepSeek WeChat group. Agentic coding benchmarks jumped sharply from the preview version: DeepSWE from 12.8 to 62.7, CyberGym from 52.7 to 83.3, and Terminal Bench 2.1 from 72.1 to 87.9 — placing V4 Pro 0813 near Fable 5 territory at $0.87/M output tokens, roughly 1/60th the cost of Fable 5, according to analysis by DigitalApplied. DeepSeek has also warned that a “significant” price increase is coming, with no date or new rate published.
What Changed: V4 Pro 0813 vs. the April Preview
DeepSeek V4 Pro 0813 is not a new product — it is the GA date-stamped release of the same MoE architecture first published in preview four months ago, with updated weights producing substantially different benchmark results. DeepSWE, which measures agentic software-engineering capability, rose from 12.8 to 62.7. CyberGym, which tests agentic security-task performance, rose from 52.7 to 83.3. Terminal Bench 2.1, an agentic coding benchmark run in a headless terminal environment, rose from 72.1 to 87.9. These figures originated from the DeepSeek WeChat group and were subsequently reported by Simon Willison at simonwillison.net and circulated on Reddit’s r/LocalLLaMA; they are preliminary, not publications from DeepSeek directly.
DeepSeek V4 Pro is a 1.6 trillion-total-parameter model with approximately 49 billion active parameters per forward pass (mixture-of-experts architecture), a 1 million-token context window, and up to 384K output tokens — confirmed via the OpenRouter listing at openrouter.ai/deepseek/deepseek-v4-pro-0813.
V4 Pro 0813 is a different model from DeepSeek V4-Flash, the smaller, faster sibling optimized for cost-per-task on bounded workloads. V4 Pro is the full MoE flagship with the larger context window and the higher benchmark scores.
Pricing: $0.87/M Output — with a Significant Hike on the Way
API pricing for DeepSeek V4 Pro 0813 is $0.435/M input tokens (cache miss), $0.003625/M input tokens (cache hit), and $0.87/M output tokens. The cache-hit rate makes repeat-context workloads — coding agents that re-read the same codebase headers on every call — exceptionally cheap: 10 million cached input tokens per month costs $36.25 at that rate.
DeepSeek has warned that a price increase is imminent. “DeepSeek says it plans to raise API prices in the near future and expects the increase to be significant, though it has not published the new rates or an effective date,” WCCFTech reported on August 12, 2026. The company has historically priced its API below cost as a market-penetration strategy; the acknowledgment of a coming hike signals the end of that subsidized phase.
Open Weights: April Preview Confirmed — 0813 Weights Unconfirmed
The April preview version of DeepSeek V4 Pro is available on HuggingFace at deepseek-ai/DeepSeek-V4-Pro. As of August 13, 2026, the 0813 GA weights have not been confirmed as published; API access via OpenRouter is the only confirmed path to the updated model. “I had to link to OpenRouter because DeepSeek don’t have any obvious announcement page for their new model,” wrote Simon Willison on simonwillison.net.
What V4 Pro 0813 Means for Business AI Buyers
For teams evaluating AI agents for business tasks, V4 Pro 0813 resets the price-performance floor for agentic coding. Benchmark parity with Fable 5 at approximately 1/60th the cost — if these preliminary DeepSWE, CyberGym, and Terminal Bench scores hold under production conditions — represents a material cost argument for any team currently paying frontier prices for coding-agent access. The pending price hike adds a concrete time-pressure element: DeepSeek has explicitly signaled that current rates are not permanent.
For businesses evaluating AI tools for business workflows that require large-context reasoning or multi-step agentic tasks, V4 Pro 0813’s 1 million-token context window and cache-hit economics are worth benchmarking directly against existing Fable 5 or GPT-5.6 Luna spend before the increase lands. Teams with high repeat-context usage — legal document analysis, large-codebase review, extended customer-service threads — stand to gain the most from the cache-hit pricing.
DeepSeek’s chip ambitions further context the pricing strategy: the company is building its own AI chip to reduce dependence on Nvidia, which would reduce its inference cost floor over time, even as it moves away from subsidized pricing on the current generation.
Our Take
Frequently Asked Questions
What are DeepSeek V4-Pro’s new API prices after the August 16, 2026 increase?
DeepSeek V4-Pro output tokens now cost $3.96/M during peak hours and $1.98/M during off-peak hours, up from a flat $0.87/M. Input tokens also increased. V4-Flash output tokens moved to $1.32/M (peak) and $0.66/M (off-peak), up from $0.28/M flat.
Why did DeepSeek raise its API prices?
DeepSeek cited “heightened demand” as the reason for the August 16, 2026 increase. The company switched from flat-rate to peak/off-peak tiered pricing, a structure that shifts cost-sensitive workloads to less-congested hours rather than continuing the below-cost subsidized rates it had used as a market-entry strategy.
How does DeepSeek V4-Pro compare to Fable 5 in terms of cost after the price hike?
At peak pricing ($3.96/M output), DeepSeek V4-Pro costs roughly 1/6th the price of Fable 5. At off-peak ($1.98/M), it is approximately 1/12th. Before the hike, it was 1/60th. The price-performance advantage remains significant, but the gap has narrowed.
What is the DeepSeek V4-Pro 0813 model?
DeepSeek V4-Pro 0813 is the general-availability release of DeepSeek’s 1.6 trillion-parameter mixture-of-experts model, date-stamped August 12, 2026. It has approximately 49 billion active parameters per forward pass, a 1 million-token context window, and benchmark scores near Fable 5 on agentic coding tasks including DeepSWE (62.7) and CyberGym (83.3).
Is DeepSeek V4-Pro available for self-hosting?
The April preview version of DeepSeek V4-Pro is available on HuggingFace at deepseek-ai/DeepSeek-V4-Pro. As of August 2026, the 0813 GA weights have not been confirmed as publicly released; API access via OpenRouter is the confirmed path to the updated model.
Price Hike Confirmed: DeepSeek Raises API Prices Up to 1,100%
DeepSeek implemented the price increase on August 16, 2026 at 16:00 UTC, switching from flat-rate pricing to a tiered peak/off-peak model across all V4-Pro and V4-Flash API tiers. DeepSeek cited “heightened demand” and said tiered pricing shifts workloads to less-congested hours — a structural change, not a temporary surcharge.
The new output-token rates for V4-Pro are $3.96/M during peak hours and $1.98/M during off-peak hours, up from a flat $0.87/M — a 355% to 455% increase depending on when you call the API. V4-Flash output tokens moved to $1.32/M (peak) and $0.66/M (off-peak), up from $0.28/M flat. Input-token rates rose in parallel; the highest individual tier-to-tier jump exceeds 1,100%, according to reporting by Crypto Briefing, Forbes, and Quartz.
For teams that built cost models on DeepSeek’s flat rates, the impact is immediate: a workload running 10 million V4-Pro output tokens per month at peak now costs $39,600 instead of $8,700. Off-peak scheduling — routing non-latency-sensitive jobs to less-congested windows — brings that down to $19,800 at $1.98/M.
The “significant hike” that DeepSeek had signaled in early August has landed. Teams evaluating AI agents for business tasks at scale should re-benchmark total cost of ownership at the new rates before committing to V4-Pro as a primary inference provider.
For Context — earlier witho2.com coverage of DeepSeek and related AI agents: DeepSeek V4-Flash beats its own flagship on agent benchmarks · DeepSeek’s chip ambitions. Related: open-weight AI agents entering the market alongside proprietary alternatives.

