OpenAI adds cheaper Sol and Luna models to GPT-6
OpenAI has expanded the GPT-6 family with GPT-6 Sol and Luna, two smaller models designed for workloads that balance capability, latency, and cost. They join the flagship GPT-6 Astra, released earlier this month. OpenAI says Sol and Luna retain many of Astra’s gains while costing 50% less than the promotional prices of their GPT-5.6 counterparts.
Sol occupies the mid-tier for coding, computer use, and tool-driven workflows. Luna targets high-volume tasks where low per-request cost matters most. Astra remains OpenAI’s most capable option for difficult professional work. All three use related training methods, allowing OpenAI to distribute improvements in factuality, coding, tool use, and alignment across different price tiers.
Prices fall by half
The cached-input figures reflect OpenAI’s 90% discount for cache hits. A workload’s final cost depends on its mix of input, cached input, output, and reasoning effort. Long-running agents can benefit disproportionately because they often resend the same system prompt, tool definitions, and conversation history across many model calls.
Luna competes with small models such as Google’s Gemini Flash and Anthropic’s Haiku. Sol sits in the middle tier, where coding agents and business automations need stronger reasoning without paying flagship rates on every step.
Sol’s case rests on cost per task
OpenAI emphasizes completed-task cost rather than raw benchmark scores. Its comparisons use effort settings such as medium, max, and xhigh, which allocate more inference-time computation to a request. These settings differ across providers, so similarly named levels are not directly equivalent.
OpenAI also tested factuality using real ChatGPT conversations in which users had flagged errors. The company says Sol cuts its predecessor’s mistake rate by roughly half. At higher reasoning settings, Luna approaches GPT-5.6 Sol’s factual reliability at about one-hundredth of the cost.
These comparisons depend on prompts, tool configuration, reasoning settings, and completion length. Production evaluations should measure task success, latency, and total cost against the applications and data a team actually uses.
Caching survives mid-session changes
Prompt caching stores a reusable prefix of a request, such as system instructions and tool schemas, so the model does not process those tokens from scratch on every call. OpenAI has improved default cache-hit rates while retaining the 90% discount on cached input reads.
Three changes target agent loops
- Developers can change reasoning effort or enable and disable tools during a conversation without invalidating the existing cache.
- Explicit cache breakpoints define where a reusable prefix ends, providing finer control over which content remains cached.
- A caching dashboard and diagnostics tool show where requests miss the cache.
OpenAI cites GitHub’s use of its models as an example. According to the company, the caching changes reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests, which helped GitHub Copilot respond faster.
Less narration in coding sessions
OpenAI says Sol and Luna inherit Astra’s revised response style, with fewer preambles, less repetition of the prompt, and shorter answers. The change is particularly relevant in coding sessions, where unnecessary narration consumes output tokens and adds latency without advancing the task.
Both models also improved over their GPT-5.6 counterparts in OpenAI’s internal red-team evaluations. The company reports lower rates of coding deception, which occurs when an agent claims to have completed changes or tests that it did not perform. OpenAI provides additional results in the system card.
The rollout varies by product
Production economics favor model routing
For teams deploying agents at scale, the release changes two cost levers at once: token prices and the amount of context that requires fresh processing. Preserving the cache when an agent changes tools or reasoning effort supports workflows that begin with a low-cost configuration and allocate more computation only when a difficult step requires it.
Sol is the model to benchmark for recurring agent workloads that need substantial reasoning but cannot justify flagship pricing on every call. Luna offers a cheaper route for simpler tasks, and Astra remains available for cases where maximizing success rate outweighs cost and latency.