AI IndustryModel Pricing & Benchmarks

DeepSeek V4 Pro vs Grok 4.6: Agent Benchmarks, Price Cuts, and the Peak-Pricing Shift

DeepSeek V4 Pro posted massive Terminal-Bench and CyberGym gains in the same week Grok 4.6 undercut frontier pricing by up to 88%, but DeepSeek's quiet move to peak/off-peak token pricing means the real comparison is total agent-workload cost, not raw benchmark scores.

6G-AI Editorial TeamAug 18, 20267 min read
Share:

A Week That Repriced the Frontier

On August 13, 2026, three frontier-class models landed within hours of each other: Grok 4.6, DeepSeek V4 Pro 0813, and Qwen3.8. Two days later, DeepSeek made a quieter but arguably more consequential move: starting August 16, its V4-Flash and V4-Pro lines began charging different prices for output tokens and cache hits depending on the UTC hour. Some cache-hit prices rose by more than ten times.

The result is that the obvious headline comparison, DeepSeek V4 Pro vs Grok 4.6 on benchmark scores, is now the wrong question. Both models are competitive on paper. The real question for anyone running agent workloads is what a full task costs once you account for token prices, cache behavior, retry loops, and the hours of the day your agents actually run. This article walks through the verified numbers from this week and then builds the cost comparison that matters.

What DeepSeek V4 Pro 0813 Actually Improved

According to the official figures circulated at launch, DeepSeek V4 Pro 0813 delivered some of the largest single-version agent benchmark jumps we have seen this cycle:

  • Terminal-Bench 2.1: from 72.1% to 87.9%
  • CyberGym: from 52.7% to 83.3%
  • DeepSWE: from 12.8% to 62.7%
  • Context window: 1 million tokens
  • List pricing: $0.435 per million input tokens, $0.87 per million output tokens

The DeepSWE jump, nearly a fivefold improvement, is the kind of number that suggests the previous version had a harness or formatting problem rather than an intelligence ceiling, a pattern we explored in our coverage of why the next agent battlefield is the harness, not the model. Terminal-Bench and CyberGym gains of 15 to 30 points in one revision are more straightforwardly impressive, and they position V4 Pro as a credible low-cost engine for terminal agents and security-flavored coding loops, a space where mid-tier models have been steadily absorbing agent workloads all year.

One caveat, also raised in the original report of these numbers: every vendor now cherry-picks the single benchmark that tells its best story. When each model markets a different leaderboard, time-to-completion, rework rate, and error type on your own tasks are more comparable metrics than any published score.

Grok 4.6 and the 88% Undercut

Grok 4.6's launch week narrative was not a benchmark record but a price attack. Investor Gavin Baker argued on X that Grok 4.6 performs close to Fable 5 Max while costing roughly 80% less on input and 88% less on output. Baker was explicit that this is a personal-experience judgment, not a unified evaluation, and it should be read that way. But the strategic point underneath his take is sound: when frontier scores cluster in a narrow band, price, response speed, rate limits, and how many turns a long task can sustain start to matter more than who holds the nominal top spot.

This is the same commodity-curve logic we traced when Claude Opus 5 arrived at half price and when Kimi K3's open-weight release repriced intelligence. Closed-source vendors are no longer selling raw intelligence alone; they are selling less rework, less waiting, and a more predictable delivery expectation. An 88% output-price cut is a direct attempt to break that premium.

The Quiet Bombshell: Peak and Off-Peak Token Pricing

The most structurally important news of the week got the least attention in the launch-day noise. Beginning August 16, DeepSeek started pricing V4-Flash and V4-Pro output tokens and cache hits differently by UTC time window, with some cache-hit prices rising more than tenfold.

Why does this matter so much? Because near-free cache hits were, in effect, a hidden subsidy for exactly the workloads this article is about. Long conversations, repeated queries, and agent tool loops, where the model re-reads a large stable context on every turn, all leaned heavily on cheap caching. An agent that makes forty tool calls against a 200,000-token context was paying a small fraction of what the raw price sheet implied. Once cache pricing scales with real datacenter load, long-horizon tasks are exposed to the physical capacity of the machine room for the first time.

Intelligence is becoming an actual utility, not just a metaphor about being as cheap as electricity. The cheapness still exists, but it has moved to off-peak hours, to local models, and to teams that know how to schedule batch jobs. This also reframes the open-versus-closed trade we discussed in Epoch AI's four-month-gap framing: self-hosting is no longer just a privacy play, it is a hedge against time-of-day pricing.

DeepSeek V4 Pro vs Grok 4.6: The Cost Comparison That Matters

So how do the two models actually compare for a buyer this week? Stack the verified facts side by side.

On benchmarks: DeepSeek publishes hard agent numbers showing dramatic improvement (87.9% Terminal-Bench 2.1, 83.3% CyberGym). Grok 4.6's claim is proximity to a top-tier model at a fraction of the price, based on individual experience rather than a standardized eval. Neither claim is directly comparable to the other, which is precisely the problem with score-first shopping.

On sticker price: DeepSeek V4 Pro lists at $0.435 per million input and $0.87 per million output, among the lowest frontier-adjacent prices on the market. Grok 4.6's claimed advantage is 80 to 88% below Fable 5 Max, which still leaves it well above DeepSeek's list prices in absolute terms, but competitive within the premium tier it is attacking.

On total agent-workload cost: this is where the comparison inverts. DeepSeek's sticker price is now a schedule, not a number. The same agent run at a peak UTC hour costs more than the identical run at an off-peak hour, and cache-heavy loops cost more than they did last week under either schedule. Grok 4.6, by contrast, is competing on flat, predictable cheapness. For a team running continuous agents around the clock, DeepSeek may still win on pure arithmetic if jobs can be shifted to off-peak windows. For a team running latency-sensitive interactive agents during business hours, the effective gap between the two models is far narrower than the list prices suggest, and Grok's speed and quota characteristics, highlighted in Baker's take, start to dominate.

The honest summary: DeepSeek V4 Pro offers the higher verified agent ceiling at the lower sticker price, with new time-based volatility attached. Grok 4.6 offers a claimed near-frontier experience with aggressive, predictable pricing and no time-of-day math. Neither is a universal winner, and anyone telling you otherwise is selling something.

Benchmarks Are Marketing; Failure Modes Are the Product

The third model released that day, Qwen3.8, barely figures in most headlines, but its presence sharpened the right question. When three frontier models land at once with scores clustered in a similar range, "use whichever scores highest" stops being a strategy. What you are actually buying is a failure mode: a cheap model may finish in minutes but leave a bug behind; an expensive model may cost more but save a rework cycle; an open-weight model moves the errors into your own inference stack, VRAM budget, and ops team.

For an engineer fixing bugs, a few dollars of price difference per task is trivially small next to one hour of rework. This is why the Terminal-Bench and CyberGym jumps matter even to buyers who distrust benchmarks: a model that goes from 52.7% to 83.3% on a structured task suite is plausibly failing less often in ways that are expensive for you to catch. It is also why record-chasing headlines like Claude Opus 4.6's SWE-bench record need the same translation: what does a few points of benchmark mean in avoided rework hours on your codebase?

Practical Routing Advice for August 2026

Given everything above, a sensible routing posture this month looks like this:

  • Batch and overnight agent jobs: DeepSeek V4 Pro, scheduled into off-peak UTC windows. The benchmark jumps are real enough to trust for unattended terminal and security loops, and off-peak pricing restores the cost advantage that cache repricing eroded.
  • Interactive, business-hours agents: benchmark both. Grok 4.6's flat cheap pricing and reported speed make it a strong default for latency-sensitive turns; DeepSeek remains worth testing if your workload is input-heavy rather than cache-heavy.
  • Long-horizon tool loops: audit your cache-hit ratio before choosing anything. Under the new DeepSeek schedule, cache behavior is now a first-order cost variable, and teams running fully local agents, as we covered in Block's open-sourced Goose and the desktop shift by Manus, now have a fresh economic argument alongside the privacy one.
  • Everything: measure time-to-done and rework rate on your own tasks for a week before committing. Vendor-picked benchmarks, including the genuinely impressive ones cited in this article, are directional signals, not contracts.

It also helps that DeepSeek V4 Pro 0813 is already in the real call chain: it appeared on OpenRouter the same night, with support for thinking mode and reasoning_effort, alongside Grok 4.6 and Qwen3.8. Model competition is now fought directly at the routing layer, where switching costs are measured in a config line, which means buyers can actually run the two-model experiment this article recommends instead of just reading about it.

What to Watch Next

Three things will decide how the DeepSeek V4 Pro vs Grok 4.6 comparison reads in a month. First, whether independent Terminal-Bench and CyberGym reproductions confirm the official jumps; the DeepSWE number in particular invites scrutiny. Second, whether Grok 4.6's price advantage survives contact with standardized evaluations, or turns out to be a great deal on a model that needs more retries, which would hand the rework-cost argument right back to DeepSeek. Third, and most important, whether peak/off-peak pricing spreads. If other providers follow DeepSeek's lead, every model comparison published today, including this one, will need a time-of-day column. The era of comparing APIs by a single price per million tokens ended on August 16, 2026. The buyers who noticed first will pay less for the same intelligence.

Share:

Related Articles