[models] · · 1 min read
OpenAI slashes GPT-5.6 Luna and Terra prices up to 80%, launches faster Sol tier
New API pricing makes Luna 80% cheaper and Terra 20% cheaper, while Sol's Fast mode delivers 2.5x speed at double the price.
By ByteBulletin Editors · Editorial Team
OpenAI today announced significant price cuts for two of its GPT-5.6 model tiers — Luna and Terra — as well as a new Fast mode for the top-tier Sol model that promises up to 2.5× faster speeds at twice the price.
The cost reductions are dramatic: Luna, the fastest and most affordable model, drops 80% to just $0.20 per million input tokens and $1.20 per million output tokens. Terra, the balanced everyday model, falls 20% to $2 per million input and $12 per million output. Sol pricing remains unchanged.
According to OpenAI, these efficiencies come from improvements across the entire stack — models, inference systems, and agentic harness. The company notes that Sol has begun autonomously rewriting production kernels and running experiments to boost token generation, contributing to a 20% reduction in end-to-end serving cost and a 15% increase in generation efficiency.
“The gains can compound,” OpenAI writes. “More capable models help our technical team find the next generation of improvements, shortening the path to better performance and lower costs.”
The new Fast mode for Sol replaces the previous Priority Processing offering and is backward compatible. For developers building agentic workflows, the tiered pricing creates real incentives to route tasks intelligently: use Sol for planning and uncertainty resolution, then hand implementation to Luna for well-defined work.
OpenAI frames this as part of its mission to make intelligence more abundant and affordable, enabled by “years of improvements in how our models are built, served, and put to work.” The pricing changes also flow through to ChatGPT Work and Codex subscription quotas, where Luna and Terra usage now consumes fewer credits.
SHARE
RELATED

[models] ·
Escha Labs ships 2-bit Qwen3.6-35B-A3B build that runs on a 24GB GPU
A 12.3GB MoE quant packs a 35B model onto consumer cards, with quality within noise of FP8 on most benchmarks.

[models] ·
ByteDance Is Training a 10 Trillion-Parameter Model, Leaning Into Scale to Chase Anthropic
The TikTok parent is reportedly pre-training a massive model that could rival Anthropic's Mythos 5, signaling a new phase in the global AI race.

[models] ·
Alibaba’s Qwen3.8-Max goes open-weight, claiming Claude-rivaling performance
Alibaba releases its largest open-weight model yet, Qwen3.8-Max, claiming it rivals Anthropic’s Claude Fable 5, with weights due next week.
