GPT-5.6 Luna: Free and Go Access, Fast Responses, and Low-Cost AI Work
GPT-5.6 Luna is OpenAI's cheapest current model, and after an 80% price cut three weeks into its life it became one of the most-used models on the open market.
........
Luna launched July 9, 2026 at $1 per million input tokens and $6 output; on July 30 OpenAI cut that by 80% to $0.20 and $1.20.
On August 6, 2026, OpenAI announced Luna would become the default model for ChatGPT Free and Go accounts, with unlimited text chats.
It shares the GPT-5.6 family's roughly 1.05 million token context window and 128,000 token maximum output.
It currently sits third on OpenRouter's usage rankings at 9.52 trillion tokens processed in a week, up 129%.
Replit's Free Mode for Core and Pro subscribers runs on Luna, the largest third-party commitment to the tier so far.
··········
WHERE LUNA SITS.
Luna is the smallest and cheapest of the three GPT-5.6 tiers, below Terra and well below the flagship Sol.
OpenAI's own framing describes it as fast and affordable, suited to high-volume, latency-sensitive work: chat, classification, and lightweight agentic workflows.
The naming scheme is deliberate. In the GPT-5.6 system the number marks the generation while Sol, Terra, and Luna mark durable capability tiers that can advance on their own schedule — which means a future Luna will still be the cheap tier, just a better one.
··········
THE PRICE CUT.
Luna's economics changed dramatically three weeks after launch.
At general availability on July 9, 2026, Luna was priced at $1 per million input tokens and $6 per million output tokens.
On July 30, 2026, OpenAI cut Luna's price by 80%, to $0.20 per million input tokens and $1.20 per million output tokens. Terra was cut 20% in the same announcement.
OpenAI framed the cut as passing along efficiency gains, having shared the day before how GPT-5.6 had helped make itself cheaper to run.
The cut also applied to how usage counts against paid subscriptions in Codex and ChatGPT Work, so subscription quotas stretch further on Luna than they did before.
··········
Luna pricing before and after July 30, 2026
At launch (July 9) | After July 30 cut | |
|---|---|---|
Input per million tokens | $1.00 | $0.20 |
Output per million tokens | $6.00 | $1.20 |
Change | — | -80% |
··········
HOW LUNA COMPARES INSIDE ITS OWN FAMILY.
The gap between the three tiers is large enough that tier choice, not prompt optimization, is usually the biggest lever on an API bill.
Terra sits at $2 input and $12 output per million tokens after its own cut, making it ten times Luna's rate.
Sol, the flagship, has moved as well: as of late August 2026 it was listed at $4 input and $20 output under promotional pricing, down from $5 and $30, with OpenAI stating that promotional rate is available at least through November 21, 2026 — a floor rather than an announced end date.
All three share the same roughly 1.05 million token context window and 128,000 token output ceiling, so the tier decision is about capability and cost rather than about how much they can read.
··········
GPT-5.6 tier pricing per million tokens
Tier | Input | Output | Multiple of Luna |
|---|---|---|---|
Luna | $0.20 | $1.20 | 1x |
Terra | $2.00 | $12.00 | 10x |
Sol (promotional) | $4.00 | $20.00 | ~18x |
··········
THE LONG-CONTEXT SURCHARGE.
One piece of fine print moves real bills and is easy to miss.
Prompts whose input exceeds roughly 272,000 tokens are billed at twice the input rate and 1.5 times the output rate — and the surcharge applies to the entire request, not just to the portion above the threshold.
Cache reads bill at a tenth of the input rate, and cache writes at 1.25 times the uncached input rate, with a 30-minute minimum cache life.
For a model whose entire appeal is cost, crossing that 272,000 token line without noticing doubles the input bill on every affected request.
··········
LUNA IN CHATGPT.
On August 6, 2026, OpenAI pushed the GPT-5.6 tiers into the consumer product, with Luna taking the entry position.
Per that announcement, Luna became the default model for Free and Go accounts, with unlimited text chats, while Sol was updated for Plus and Pro.
File uploads, image generation, voice, data analysis, and other tools remain subject to their own separate limits regardless of the unlimited text chat allowance.
There is a wrinkle that has caused confusion since general availability: Terra and Luna are not selectable models in standard ChatGPT conversations. Where they are individually selectable is in ChatGPT Work and in Codex, depending on plan. That is a surfacing decision rather than an entitlement problem, but the announcements did not make it obvious, and general availability day produced a run of complaints from paying users who could not find the tiers they had read about.
Coverage of ChatGPT plan tiers has been inconsistent on exactly which models appear on which consumer plan, and rollouts vary by account and region, so treat any model list as a snapshot rather than a guarantee.
··········
WHAT LUNA CAN ACTUALLY DO.
The interesting claim about Luna is not that it is cheap but that it is capable enough to run agent loops rather than single calls.
OpenAI published customer accounts describing the shift. One reported that Luna moved them from a single structured-output call to a full tool-calling agent loop, raising prompt-cache reuse from 24% to 90%.
The same account reported that across thousands of production calls, Luna handled 2.2 times more context with 8.5 times fewer output tokens, at 87% lower cost than GPT-5.4 mini.
Another customer said Terra and Luna led on cost-efficiency across their internal coding benchmarks, and that Luna had become their default model for background agent automations.
On published benchmarks, Luna trails the previous flagship GPT-5.5 on coding, though narrowly. On at least one professional clinical evaluation, Luna edges ahead of GPT-5.5.
··········
ADOPTION IN THE OPEN MARKET.
Usage data from OpenRouter shows the price cut translating into traffic.
As of usage data through September 1, 2026, GPT-5.6 Luna ranked third on OpenRouter's leaderboard with 9.52 trillion tokens processed in the trailing week, up 129% week over week.
That places it directly among the cheap high-volume models it was priced to compete with — DeepSeek V4 Flash and GLM 5.3 Flash occupy the two positions above it.
On OpenRouter, Luna is served by three providers: OpenAI directly, Azure in the US, and Amazon Bedrock in the US, with automatic failover between them.
The largest single third-party commitment came on August 19, 2026, when OpenAI and Replit announced that Replit's new Free Mode for Core and Pro subscribers runs on Luna.
··········
WHEN LUNA IS THE RIGHT CHOICE.
Luna's case is strongest wherever the same operation runs many thousands of times.
Classification, routing, summarization, extraction, background automations, and first-pass agent work are the natural fits — tasks where the cost of a single wrong answer is low and easily caught, and where volume makes per-token rates decisive.
The framing that holds up best is per task: what is the cost of being wrong here, and what is the cheapest model that clears that bar.
For high-stakes, low-volume decisions — an architecture call, a contract review, an evaluation judge — the price of a mistake dwarfs the token cost and the reasoning tiers are worth their premium.
The common production pattern is routing most traffic to Luna and escalating only the cases that need it, which is precisely the structure the tier naming was designed to make easy to reason about.
··········
·····
FOLLOW US FOR MORE.
·····
·····
DATA STUDIOS
·····



