Sunday, August 16, 2026

Today’s Edition

AI Intel Report

MARKETS

Frontier Models

DeepSeek-V4-Pro GA Delivers 1.7T Agentic Model Under MIT License With Pricing Tiers

The August 13 rollout adds Responses API compatibility and flexible reasoning controls while shifting to demand-based rates that halve during off-peak hours starting August 16.

4 MIN READ
Inside a vast temperature-controlled data center facility multiple long parallel aisles stretch into the distance lined with identical tall black server racks each rack featuring evenly spaced vertical slots holding dense compute blades with small blue indicator LEDs illuminated in steady patterns and thick bundles of fiber optic cables routed in overhead trays and neatly secured along the sides with velcro ties the floor consists of raised white perforated tiles allowing cool air to flow upward from underfloor plenums while the ceiling holds exposed HVAC ducts and fire suppression pipes a group of three anonymous technicians wearing light gray coveralls and safety glasses stand with backs facing the viewer one technician holds a tablet device and points toward an open rack door revealing internal components another technician adjusts a cable connection while the third observes a nearby control panel area the environment includes visible cooling fans spinning inside the racks condensation on chilled water pipes running along the walls and faint mist from humidification systems the overall scene conveys operational activity during reduced demand periods with some rack sections powered down showing darker unlit sections contrasted against active glowing areas symbolizing flexible resource allocation and off-peak rate adjustments the technicians appear engaged in routine monitoring and configuration tasks related to large scale AI inference hardware deployment with no visible screens displaying any readable content or branding the background shows additional racks receding into perspective with consistent spacing and identical hardware configurations emphasizing scale and uniformity of the installation environment subtle reflections on polished metal surfaces from overhead lighting fixtures create depth and the air appears slightly hazy from particulate filters in the ventilation system overall composition focuses on the hardware infrastructure and human operators interacting directly with the physical systems in a realistic industrial setting without any superimposed elements or distractions
Illustration: AI Intel Report

DeepSeek-V4-Pro is a 1.7 trillion parameter frontier model released by DeepSeek under the MIT license that delivers premium agentic performance with general availability on app, web, and API.

The general availability rollout on August 13, 2026, extends the model to production users across mobile app, web dashboard, and programmatic API endpoints while keeping the model identifier deepseek-v4-pro unchanged for existing integrations.

What background led to the V4-Pro general availability announcement?

DeepSeek previously released earlier variants including DeepSeek-V4-Flash-0731 that focused on speed optimizations. The V4-Pro line targets higher capability thresholds for agentic tasks that require extended reasoning chains and tool use. Company documentation indicates the new release consolidates prior experimental features into a stable offering suitable for enterprise workloads.

The MIT license on the 1.7T parameter weights hosted at Hugging Face allows downstream modification and commercial redistribution without restrictive terms. This licensing choice aligns with broader industry trends toward open weights for large-scale models while maintaining API revenue through hosted inference.

What new capabilities arrive with the V4-Pro general availability?

The model introduces native compatibility with the OpenAI Responses API format that enables direct migration for applications already built around Codex-style completions. Developers can activate one-click setup within supported environments. Three discrete thinking effort settings allow runtime adjustment: low for routine queries, high for standard agent loops, and max for multi-step planning scenarios.

Benchmark results released alongside the announcement include HLE scores of 42.7 without tools rising to 60.0 when tools are available. Terminal Bench 2.1 reaches 87.9 under the reported evaluation protocol. These figures position the model among leading open-weight systems for agent-oriented workloads.

What technical specifications define DeepSeek-V4-Pro-0813?

The architecture supports a 1 million token context window that accommodates large codebases or extended conversation histories. Concurrency is capped at 500 simultaneous requests for the V4-Pro tier to maintain service stability. The model name for API calls remains deepseek-v4-pro following the GA transition.

Reported agent capability benchmarks for DeepSeek-V4-Pro-0813
BenchmarkScore Without ToolsScore With Tools
HLE42.760.0
Terminal Bench 2.187.9N/A

How does the new peak and off-peak pricing structure operate?

Effective at 16:00 UTC on August 16, 2026, the API adopts a two-tier rate schedule. Off-peak periods receive a 50 percent discount relative to peak rates. The peak input price for cache misses on DeepSeek-V4-Pro-0813 stands at 1.32 dollars per million tokens according to the published pricing table.

The adjustment aims to balance compute demand across daily cycles. Users running continuous agent fleets can schedule non-urgent tasks into off-peak windows to reduce operational costs while preserving performance during business hours.

What market implications follow from the open weights and pricing changes?

Availability of the full 1.7T weights under MIT terms lowers barriers for organizations that prefer on-premises or private-cloud deployments. At the same time the hosted API with differentiated pricing provides a consumption-based alternative for teams without dedicated infrastructure.

Stakeholders in regulated sectors gain an additional option that combines high benchmark scores with transparent licensing. The concurrency limit and context window support concurrent multi-agent systems without immediate scaling constraints for mid-size deployments.

What reactions have emerged from industry observers?

We’re launching DeepSeek-V4-Pro today! 🚀 Major Agent upgrades with strong production gains! Flexible reasoning effort for V4-Pro & V4-Flash: low for simple tasks, high for daily Agent workflows, max for complex tasks. Native OpenAI Responses API support, optimized for Codex with one-click setup.DeepSeek

Analysts note that the combination of open weights and API pricing tiers creates dual pathways for adoption. Enterprises can prototype on the hosted service before deciding on self-hosted implementations using the released checkpoints.

What developments are expected next for the DeepSeek model family?

Future updates may extend the thinking effort controls to additional model variants. Continued benchmark reporting will clarify performance on new agent evaluation suites. The pricing schedule will be monitored for adjustments based on observed utilization patterns after the August 16 transition.

  1. Review current API usage patterns to identify off-peak scheduling opportunities.
  2. Test Responses API integration using the one-click Codex setup path.
  3. Evaluate benchmark results against internal agent workloads before full migration.
  4. Monitor the official changelog for any post-GA parameter or pricing refinements.

Frequently asked

When does the peak and off-peak pricing take effect?

New prices apply at 16:00 UTC on August 16, 2026. Off-peak rates equal 50 percent of the corresponding peak rates for the same token category.

Does the model name change for existing API users?

The model identifier deepseek-v4-pro remains unchanged after the general availability rollout on August 13, 2026.

Sources

  1. DeepSeek — The GA release of DeepSeek-V4-Pro has been rolled out on the APP, Web, and API.
  2. Hugging Face — DeepSeek-V4-Pro-0813 is the official release with 1.7T params under MIT License and benchmarks matching GA announcement.
  3. DeepSeek — DeepSeek-V4-Pro-0813 pricing table with peak/off-peak rates; off-peak half of peak; context 1M; concurrency 500.
  4. DeepSeek — The GA release of DeepSeek-V4-Pro has been rolled out on the APP, Web, and API. ... API Pricing Adjustment ... new prices will take effect at 16:00 (UTC Time) on August 16, 2026.