DeepSeek’s new API prices took effect at 16:00 UTC on August 16, three days after the company moved DeepSeek-V4-Pro to general availability on August 13. The headline number is V4-Pro output at $3.96 per million tokens during peak hours, up from a flat $0.87. The number that matters more for anyone running a bill is the off-peak rate, $1.98, because that is the floor now: every cell in the rate card went up, including the cheapest hour of the cheapest model.
The pricing gets more space here than the model, because since April DeepSeek’s list price has been the reference point every other lab had to explain itself against, and it just moved.
What shipped on August 13
DeepSeek-V4-Pro-0813 replaced the April preview behind the same deepseek-v4-pro model name. The 1M-token context and 384K maximum output are unchanged, and so is the architecture, still 1.6 trillion total parameters with 49 billion active per token. What changed is post-training, the same treatment that let the 284B V4-Flash beat this model’s preview two weeks earlier.
DeepSeek’s own table puts Terminal Bench 2.1 at 87.9 (72.1 for the preview), DeepSWE at 62.7 (12.8), CyberGym at 83.3 (52.7), NL2Repo at 61.5 and Toolathlon-Verified at 74.1. Those are DeepSeek’s numbers, produced inside DeepSeek’s own agent harness, and the same table shows V4-Pro losing to Kimi K3 on Terminal Bench (88.3), DeepSWE (67.5) and Toolathlon (76.5). We cover the harness and what it means for reading these scores separately; the short version is that a self-graded jump from 12.8 to 62.7 wants an outside reproduction before it goes into a procurement deck.
The outside number that exists is coarser. Artificial Analysis scores V4-Pro-0813 at 53 on its Intelligence Index v4.1.1, one point above the July V4-Flash, seven behind Kimi K3 at 60 and ten behind Claude Opus 5 at 63. That is a solid open-weight model in third place, not a frontier model, and it is the capability the new prices are attached to.
The release also added three thinking-effort levels (low, high, max) for both V4 models and native Responses API support for V4-Pro, which had been Flash-only.
The rate card, before and after
The old prices below are what DeepSeek’s pricing page showed through August 13; the new ones are what it shows now. All figures are USD per million tokens.
| Model and token type | Before | Off-peak now | Peak now |
|---|---|---|---|
| V4-Pro input, cache hit | $0.003625 | $0.022 | $0.044 |
| V4-Pro input, cache miss | $0.435 | $0.66 | $1.32 |
| V4-Pro output | $0.87 | $1.98 | $3.96 |
| V4-Flash input, cache hit | $0.0028 | $0.007 | $0.014 |
| V4-Flash input, cache miss | $0.14 | $0.22 | $0.44 |
| V4-Flash output | $0.28 | $0.66 | $1.32 |
Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, seven hours a day, which is 9:00 to noon and 14:00 to 18:00 in Beijing. Everything else is off-peak, and off-peak is exactly half of peak on every line.
Run the divisions and the multipliers spread from 1.5x (V4-Pro cache-miss input, off-peak) to 12.1x (V4-Pro cache-hit input, peak). Output, the line most people quote, is 2.3x off-peak and 4.6x at peak for Pro, and 2.4x and 4.7x for Flash. Cache hits saw the biggest proportional increase on both models: the cache discount on V4-Pro shrank from roughly 1/120th of the cache-miss price to 1/30th.
That spread is why the coverage disagreed with itself. Caixin’s “50% to 1,100%” is the full range from the cheapest cell to the most expensive one; in DeepSeek’s yuan pricing (3 yuan to 4.5 for off-peak input, 0.025 to 0.3 for peak cache hits) the endpoints work out to exactly 50% and 1,100%. Engadget’s “four times” is V4-Pro output at peak. InfoWorld’s “more than 10x” is the cache-hit line at peak. All of those figures are in the table, and none of them is “the” increase, because there isn’t one. There are twelve, and which one applies to you depends on your model, your cache-hit rate, and what time of day your traffic runs.
The discount was labeled a promotion from day one
The more useful history is on DeepSeek’s own pricing page, where the archived footnotes lay out the sequence without needing anyone’s commentary.
V4-Pro launched on April 24 with a list price of $1.74 input and $3.48 output, and a 75% promotional discount through May 31 that produced the $0.435 and $0.87 rates most people think of as DeepSeek’s price. On May 22 the footnote changed to say the discounted rate would become the official price after the promotion ended. Then, by August 1, a new footnote: peak/off-peak pricing was coming, peak would be double the regular price, and the effective date would follow. By August 6 that footnote had been swapped for a blunter one, “a significant increase expected, please plan your usage accordingly.” On August 13 the table arrived, and the off-peak rate that the August 1 note implied would be the old price turned out to be 1.5x to 6x the old price.
So the “permanent” $0.435/$0.87 lasted from May 22 to August 16, under three months. The new peak input rate of $1.32 is still below the April list price; the new peak output rate of $3.96 is above it. DeepSeek’s stated reason, in the changelog, is “to allocate resources more reasonably” by pushing users to schedule work off-peak. Reuters and Caixin, which broke the story on August 13 and 14, reported no fuller explanation, and InfoWorld’s framing of compute strain is analyst inference rather than a company statement.
DeepSeek has priced by time of day before, when it discounted V3 and R1 off-peak in early 2025, and it is standard practice for infrastructure businesses generally. The direction is what changed. The DeepSeek pricing moves that made headlines were cuts, and cuts from a lab that had never published its unit economics were easy to read as a structural cost advantage. A lab that raises prices within three months of making a discount permanent is telling you the discount was a customer-acquisition budget, and the budget ran out at roughly the point where the product got good enough to charge for.
What it does to a cost model
For a US-based operator, the practical increase is the off-peak column. Both peak windows fall between 9 p.m. and 6 a.m. Eastern (6 p.m. to 3 a.m. Pacific), so a workload that runs during US business hours never sees the peak rate. Overnight batch jobs scheduled at 2 a.m. Eastern do, and should be moved.
Take an agent workload that consumes 10 million input tokens a day at a 70% cache-hit rate and produces 2 million output tokens, on V4-Pro. Under the old prices that was about $3.07 a day. Off-peak now it is about $6.09, and at peak about $12.19. Call it 2x during US hours and 4x if you let it drift into the Beijing morning. The same workload on Flash goes from about $1.00 to $2.03 off-peak.
The comparison that shaped last quarter’s build decisions has moved but not reversed. On output at peak, V4-Pro is now roughly a sixth of Claude Opus 5’s $25 and about a quarter of Kimi K3’s $15 hosted rate; before August 16 it was about a twenty-ninth and a seventeenth. DeepSeek is still the cheap option among models scoring in the 50s and 60s, though no longer so cheap that price alone ends the evaluation. Artificial Analysis’s cost-to-run-the-index figure for V4-Pro ($604.51 across 130 million output tokens) will be worth rechecking as the new rates flow through, since verbosity and per-token price multiply.
The planning lesson is the one we drew in the production cost breakdown in June: model list price is one input to a bill that also depends on cache behavior, retry rates, and routing, and it is the input the vendor controls unilaterally. A cost projection that assumed DeepSeek’s rates were a durable feature of the market, rather than a number a private company can double with three days’ notice, was built on the wrong assumption. That goes for Qwen, Moonshot and Zhipu as much as for DeepSeek, and for Anthropic and OpenAI, whose price cuts this summer were as unilateral as this increase. Treat every provider’s price as a spot rate: keep the abstraction layer that lets you move traffic and the eval set that tells you whether the cheaper model is good enough for your tasks, and re-run both when a rate card changes.
What to watch
The obvious thing is whether the peak rate holds. DeepSeek’s only stated rationale is resource allocation, and if the split is really about smoothing load, peak prices could come down as usage redistributes or as more hardware comes online. If they stay put through the autumn while off-peak creeps up, the load-balancing framing was a softer way of announcing a general increase.
Whether Alibaba, Moonshot and Zhipu follow matters as much. All three priced against DeepSeek through the spring. Kimi K3’s hosted API was already 17x V4-Pro’s old output rate; if Moonshot holds while DeepSeek climbs, the gap between the two most-cited Chinese labs narrows to under 4x at peak, and “Chinese-lab pricing” stops being a single number anyone can put in a spreadsheet. If they raise instead, the price floor for open-weight frontier models has moved for good, and the self-hosting math that never quite worked at April prices deserves another look.
And an independent Terminal Bench 2.1 run of V4-Pro-0813 outside DeepSeek’s harness would settle the question the price raises, because the 87.9 is the figure that would justify it. If it reproduces within a few points on a neutral harness, DeepSeek is charging more for a model that got a lot better, which is ordinary. If it lands closer to the Artificial Analysis picture, a mid-50s model at 2x to 4x the old price is a different product from the one that was announced.
Sources
- Models & Pricing (live rate card, effective Aug 16 2026) — DeepSeek API Docs
- Change Log: DeepSeek-V4-Pro Update, 2026-08-13 — DeepSeek API Docs
- DeepSeek-V4-Pro GA Release — DeepSeek API Docs
- Models & Pricing, archived Aug 11 2026 (pre-increase table and footnotes) — Internet Archive
- Models & Pricing, archived May 22 2026 (75% promotion footnote) — Internet Archive
- DeepSeek V4 Pro 0813 model page — Artificial Analysis
- DeepSeek V4 Pro 0813 vs Kimi K3 comparison — Artificial Analysis
- Claude Opus 5 model page — Artificial Analysis
- DeepSeek Launches V4-Pro and Raises API Prices by as Much as 1,100% — Caixin Global
- DeepSeek raises API pricing for its V4 models — Reuters via Investing.com
- DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity — InfoWorld
- DeepSeek’s AI models are about to cost four times more — Engadget
- Deepseek ships improved V4 Pro, open-sources its agent software, and raises API prices — The Decoder
- DeepSeek V4-Pro 75% Price Cut Goes Permanent — Codersera
