Tuesday, August 11, 2026

Today’s Edition

AI Intel Report

MARKETS

Frontier Models

OpenAI Cuts GPT-5.6 Luna Prices by 80 Percent to $0.20 per Million Input Tokens

Permanent reductions driven by inference efficiency position Luna for agentic workloads while GPT-5.6 Terra drops 20 percent and Google counters with Gemini 3.6 Flash at $1.50 input pricing.

4 MIN READ
Inside a vast climate-controlled data center facility the scene shows rows of tall black server racks densely packed with graphics processing units networking cards and high-speed storage arrays all actively running inference workloads for advanced artificial intelligence systems. Anonymous technicians in neutral business casual clothing and identification badges work methodically at multiple stations one technician viewed from behind adjusts cabling on an open rack door revealing organized bundles of fiber optic lines and power distribution units another technician sits at a long metal workbench examining hardware components such as accelerator cards and cooling modules spread across the surface a third technician stands near a monitoring console reviewing system performance indicators on screens that display abstract graphs and numerical readouts without any legible text or branding. Overhead industrial air handling units and chilled water pipes run along the ceiling providing efficient thermal management while raised access flooring allows underfloor cabling and power delivery throughout the space. In the midground several wheeled carts hold spare server components and diagnostic tools emphasizing streamlined operations that reduce computational overhead. The entire environment features uniform gray and blue tones with subtle indicator lights blinking in patterns that suggest optimized resource allocation for large-scale model execution. Distant background reveals additional parallel aisles of identical racks extending into the facility depth with emergency exit signage and safety equipment mounted on walls all contributing to an atmosphere of permanent infrastructure improvements enabling lower operational expenses. The composition centers on the hardware density and human activity focused on maintenance and monitoring without any specific individuals facing the viewer or any visible corporate identifiers instead highlighting generic elements of modern AI inference infrastructure that support cost-efficient deployment of models for agentic applications and competitive positioning against rival offerings in the same category. Additional details include visible uninterruptible power supply cabinets neatly arranged along one wall bundles of ethernet cables color-coded for different network segments and ventilation grilles integrated into the rack doors all contributing to a hyper-detailed portrayal of efficiency gains in real-world AI production environments where reduced resource consumption directly translates to pricing advantages for specific model variants such as those optimized for input token processing in high-volume workloads.
Illustration: AI Intel Report

GPT-5.6 Luna is OpenAI's fastest and most affordable frontier model after an 80 percent price reduction to $0.20 per million input tokens.

OpenAI announced an 80 percent reduction in the pricing of its GPT-5.6 Luna model as part of updates to the GPT-5.6 family. The change is effective July 30, 2026, and applies permanently to the API. This adjustment follows efficiency gains achieved in inference and serving processes.

What prompted the pricing adjustments for GPT-5.6 models?

The primary driver behind the price cuts is the achievement of greater efficiency in how the models are run. OpenAI has indicated that these improvements allow for lower costs without the need for loss leading. The Luna model benefits most from this development as the fastest and most affordable option.

Prior to the update Luna pricing was higher and the reduction brings it to a level that supports broader adoption in cost sensitive scenarios. Terra also receives a reduction though smaller in scale. Sol sees no change but gains a fast mode feature for certain operations.

What are the new prices for GPT-5.6 Luna, Terra, and Sol?

The new input token price for GPT-5.6 Luna is $0.20 per million. The output token price for the same model is $1.20 per million. These figures reflect the full 80 percent reduction applied to input tokens from earlier rates.

For GPT-5.6 Terra the input price is now $2 per million tokens. The output price is $12 per million tokens. This constitutes the 20 percent reduction from earlier pricing levels.

GPT-5.6 Sol pricing has not changed from its previous levels. Users of this model can now access a fast mode that delivers up to 2.5 times the speed in certain operations without any adjustment to base rates.

Comparison of API pricing for selected frontier models
ModelInput Price ($ per million tokens)Output Price ($ per million tokens)Price Change
GPT-5.6 Luna0.201.20-80%
GPT-5.6 Terra2.0012.00-20%
GPT-5.6 SolUnchangedUnchanged0%
Gemini 3.6 Flash1.507.50N/A

How does Gemini 3.6 Flash factor into the pricing competition?

Google has introduced Gemini 3.6 Flash with input pricing set at $1.50 per million tokens. The output pricing for this model is $7.50 per million tokens. It is designed for agentic and coding workloads where token efficiency provides additional advantages.

This pricing allows Gemini 3.6 Flash to offer competitive costs for building and running agents. The model is presented as reducing the overall cost per agentic task through its design and efficiency gains.

major price cuts today: *80% drop for GPT-5.6 Luna, now $0.20 per million input tokens and $1.20 per million output *20% drop for GPT-5.6 TerraSam Altman, CEO, OpenAI

What market implications arise from these price reductions?

The lower pricing for Luna can lead to increased usage of agentic AI by lowering the barrier for frequent calls. Stakeholders such as developers and enterprises may find it more feasible to deploy larger scale agent systems.

The overall effect is an escalation in the price war for agent focused models. Both OpenAI and Google are adjusting their offerings to capture share in this growing segment of the market.

  1. Agent developers benefit from reduced costs per task when using the updated Luna pricing.
  2. Efficiency improvements provide a sustainable basis for ongoing price competitiveness.
  3. The addition of fast mode to Sol offers performance options without price increases.
  4. Competition with Gemini 3.6 Flash may lead to further optimizations by both companies.

What comes next for frontier model pricing?

Further adjustments may occur as efficiency technologies advance in the industry. The focus on agentic workloads suggests continued emphasis on cost per task metrics across providers.

Companies will likely continue to highlight how their models achieve better price performance through technical means. This trend supports wider accessibility to advanced AI capabilities for a range of applications.

Frequently asked

What is the new price for GPT-5.6 Luna input tokens?

The new price is $0.20 per million input tokens following the 80 percent cut announced by OpenAI effective July 30, 2026.

Sources

  1. OpenAI — OpenAI announced an 80% reduction in GPT-5.6 Luna pricing to $0.20 per million input tokens and $1.20 per million output tokens effective July 30, 2026, along with a 20% reduction for GPT-5.6 Terra.
  2. X — Sam Altman announced major price cuts for GPT-5.6 Luna and Terra.
  3. Google — Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens for agentic workloads.