← 返回简报

DeepSeek Ships V4 Pro as Its Flagship Model Leaves Preview

Unite.AI · 2026-08-12 19:21

AI Models & Platforms

DeepSeek Ships V4 Pro as Its Flagship Model Leaves Preview

Add Unite.AI to your preferred sources on GoogleDeepSeek has released the production version of its flagship model, DeepSeek V4 Pro, ending a preview period that ran nearly four months. The build, designated V4 Pro 0813, appeared on August 12, 2026, as the general-availability release on OpenRouter’s model page, and DeepSeek’s own API documentation now lists DeepSeek-V4-Pro-0813 as the model version behind the deepseek-v4-pro endpoint.

The API economics carry over from preview: $0.435 per million input tokens on a cache miss, $0.003625 per million on a cache hit, and $0.87 per million output tokens, with a one-million-token context window and a maximum output of 384,000 tokens. DeepSeek’s first-party pricing page confirms the 0813 version string, the pricing, and a concurrency limit of 500 for the Pro endpoint against 2,500 for Flash.

The GA caps a staged rollout DeepSeek has run in public since spring. The company previewed the V4 series on April 24, 2026, shipping open weights for both Pro and Flash under the MIT license alongside API access. On July 31, 2026, it graduated the smaller Flash model to official status and said in its change log that the Pro’s official release “will follow soon.” That follow-through is what landed with the 0813 build.

What the 0813 Build Sits On

V4 Pro is a mixture-of-experts system with 1.6 trillion total parameters and 49 billion active per token, per the model card on Hugging Face. The architecture combines two attention variants DeepSeek calls Compressed Sparse Attention and Heavily Compressed Attention, which the company says cut single-token inference compute to 27 percent and KV cache to 10 percent of what its V3.2 generation needed at the million-token setting. Both V4 models were pre-trained on more than 32 trillion tokens, with post-training that grows domain-specific experts separately and then consolidates them into one model through on-policy distillation.

Open weights for the preview builds have been downloadable since April, and the Pro repo logged more than 1.4 million downloads in the last month on Hugging Face. The card recommends a context window of at least 384,000 tokens when running the model’s maximum reasoning mode locally.

The Numbers DeepSeek Is Putting Forward

The headline evaluation claims are vendor-reported, from the model card and the company’s own harness runs. At its maximum reasoning effort, which DeepSeek labels V4-Pro-Max, the card reports:

- SWE-bench Verified: 80.6 percent resolved

- Terminal Bench 2.0: 67.9 percent accuracy

- GPQA Diamond: 90.1 percent pass@1

- Humanity’s Last Exam: 37.7 percent pass@1

- MMLU-Pro: 87.5 percent

- LiveCodeBench: 93.5 percent pass@1

- Codeforces rating: 3,206

- MRCR at one million tokens: 83.5 MMR

The card’s own comparison table places those scores against named frontier systems: V4-Pro-Max trails GPT-5.4 at xHigh effort on Terminal Bench 2.0 (67.9 to 75.1) and Gemini-3.1-Pro on Humanity’s Last Exam (37.7 to 44.4), while posting the table’s top LiveCodeBench and Apex Shortlist scores. On SWE-bench Verified it lands at 80.6, level with Gemini-3.1-Pro and a fraction behind Claude Opus 4.6 at 80.8. None of these columns has yet been replicated by an independent evaluator for the 0813 build.

The API exposes three operating modes (non-thinking, a high reasoning effort, and a max effort that the documentation describes as pushing “the boundary of model reasoning capability”) and supports the OpenAI ChatCompletions format, the Anthropic Messages format, and DeepSeek’s own Responses API, with tool calling and JSON output on both Pro and Flash.

How DeepSeek Got Here

The Pro GA completes the second half of a release strategy that put the cheaper model in front first. When V4-Flash went official on July 31, 2026, DeepSeek published agent-benchmark results showing the re-post-trained Flash build outscoring the V4-Pro-Preview on its internal coding-agent suites, a deliberate move that made the small model the default for agent workloads while the flagship stayed in preview. The 0813 build is the flagship’s answer, and it arrives with the Pro endpoint’s context window, output ceiling, and feature set unchanged in name. The change is that the preview label is gone.

The pricing posture bears watching. A notice on DeepSeek’s pricing page states the company plans “a significant increase” in overall API pricing in the near future, with specifics to come by official notice. For now, the listed V4 Pro rates hold at their preview levels, and the average price actually paid through OpenRouter sits well below the $0.435 list price, which OpenRouter attributes to caching and discounts.

DeepSeek has spent the past year building outward from the model itself: the V4 series is trained around the agentic workloads (coding assistants, multi-step automation, long-document synthesis) that its April preview announcement said already drive the company’s own in-house development. Chinese open-weight labs have been shipping flagship-class models on a near-monthly cadence, from Moonshot’s Kimi K3 to MiniMax’s agent-focused M2.7, and V4 Pro’s GA is DeepSeek’s bid to keep its flagship at the front of that pack.

What Ships Next

The open item is new weights. The Hugging Face repositories still host the April preview builds, and the model card’s download table points to those artifacts; DeepSeek has not announced a timeline for publishing 0813 weights or said whether the GA build differs from preview beyond post-training. The company’s stated cadence for the V4 line, per its change log, runs through the API first.

On the commercial side, the pricing-page notice commits DeepSeek to an official announcement of its revised price plan, with the increase applying across API services. Until that notice lands, deepseek-v4-pro continues to serve the 0813 build at the rates published on August 12, 2026 — $0.435 in, $0.87 out, per million tokens, at one million tokens of context.

本地存档正文,来自 Unite.AI,内容版权归原作者/媒体所有,仅供个人阅读存档使用。