The Impact Of Half-Price GPT‑6 Sol And Luna Models On AI Benchmark Stability
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Impact Of Half-Price GPT‑6 Sol And Luna Models On AI Benchmark Stability on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI released GPT‑6 Sol and Luna models at 50% lower prices, leading to stable but mixed performance results. Cost reductions are clear, but quality and benchmark stability are uneven, sparking industry debate.

OpenAI introduced the GPT‑6 Sol and GPT‑6 Luna models on September 22, 2026, at **half the price** of their GPT‑5.6 predecessors, aiming to democratize access to advanced AI by significantly reducing costs while maintaining competitive performance.

The new models are designed with improved caching and inference techniques, enabling OpenAI to offer them at lower prices — $2.00 per 1 million input tokens and $10.00 per 1 million output tokens for Sol, and $0.10 and $0.50 respectively for Luna. These prices represent a 50% reduction compared to GPT‑5.6, with savings passed directly to users.

Independent analysis by Artificial Analysis indicates that, despite stable composite scores across various benchmarks, the cost per task has halved, with GPT‑6 Sol at about $1.06 per task and Luna at $0.07, compared to previous models. However, some evaluation metrics, especially in knowledge work tasks, have shown regressions, attributed to changes in output presentation quality and content completeness.

Performance on hallucination reduction is notable, with Sol decreasing hallucination rates from 92% to 60% and Luna from 93% to 77%, primarily due to increased refusal rates rather than improved accuracy. This trade-off suggests a shift in model behavior, favoring fewer false answers but more cautious responses.

At a glance
reportWhen: announced September 22, 2026
The developmentOpenAI launched GPT‑6 Sol and Luna models on September 22, 2026, at half the previous prices, affecting AI benchmark scores and cost-efficiency metrics.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$86,188▲ 1.0%
Ethereum ETH$2,745▲ 0.8%
Tether USDT$0.9998▲ 0.0%
BNB BNB$788.78▲ 0.5%
XRP XRP$1.62▲ 6.7%
USDC USDC$0.9999▲ 0.0%
Solana SOL$118.07▲ 1.6%
TRON TRX$0.3437▼ 1.3%
Live data · CoinGecko · alternative.me (24h change)

GPT‑6 Sol and Luna: half the price, about the same intelligence

OpenAI’s September 22, 2026 release doesn’t raise the ceiling. It lowers the cost of everything below it, which changes what’s worth automating.

GPT‑6 Sol
$4 / $20 → $2 / $10
GPT‑6 Luna
$0.20 / $1.20 → $0.10 / $0.50

Per 1M input / output tokens. Cached input reads keep the 90% discount.

Cost per task, halved

Measured by Artificial Analysis as the weighted cost of one Intelligence Index task, at max effort.

GPT‑5.6 Sol
$1.99
GPT‑6 Sol
$1.06
GPT‑5.6 Luna
$0.18
GPT‑6 Luna
$0.07

The effort dial moves cost more than the model choice

Model and effortIntelligence IndexCost per task
GPT‑6 Sol (max)48$1.06
GPT‑6 Sol (low)34$0.13
GPT‑6 Luna (max)37$0.07
GPT‑6 Luna (low)21$0.0045
GPT‑6 Luna (non‑reasoning)18$0.01

Sol at low effort keeps about 70% of its max score for roughly an eighth of the cost, because it writes far fewer reasoning tokens. For reference, Claude Opus 5.5 leads the same index at 58.

What got better, and what got worse

Better

  • Hallucination rate on AA‑Omniscience: Sol 92% → 60%, Luna 93% → 77%
  • Coding Agent Index: Sol 57, up 2 points, at ~50% lower cost per task
  • OpenAI reports about half as many factual mistakes for Sol as its predecessor
  • Higher cache hit rates; GitHub reports over 50% fewer prompt tokens needing fresh processing

Sol gets there partly by declining more: it attempts 83% of questions vs 99%, and accuracy falls 59% → 54%.

Worse

  • GDPval‑AA v2.1: Sol down ~100 Elo, Luna down ~75
  • AA‑Briefcase v1.1: Luna down ~45 Elo
  • Coding Agent Index: Luna 41, down 2 points
  • Both models write more output tokens per task than their predecessors

Reviewers attribute the drops to weaker presentation and deliverables that omit required elements.

What to do about it

Already on GPT‑5.6 Sol or Luna? The move is mostly a price cut. Re‑test first if your output is a document someone reads, not data a system consumes.
Shelved an automation on cost? Token prices halved and the effort dial adds another order of magnitude. Re‑run the business case.
Choosing between labs? The question is no longer which model is smartest, but which clears your quality bar at the lowest cost per task.
ThorstenMeyerAI.comSources: OpenAI (pricing, vendor benchmarks) and Artificial Analysis (independent evaluation and model pages). Figures as of 23 September 2026.

Implications for Cost-Effective AI Deployment

The release of GPT‑6 Sol and Luna at half the previous prices could dramatically lower barriers for integrating AI into products and workflows, enabling more organizations to automate tasks previously deemed too costly. While cost savings are clear, mixed performance results and regressions in certain benchmarks highlight potential trade-offs in output quality and reliability, especially for knowledge-intensive tasks.

This development might accelerate AI adoption across industries, but it also raises questions about benchmark stability and whether the models’ reduced presentation quality impacts real-world applications requiring detailed, accurate outputs.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Industry Context and Prior Developments

Prior to this release, OpenAI’s models, including GPT‑5.6, had been the industry standard for high-performance AI, with pricing reflecting their advanced capabilities. The introduction of Astra, the top-tier model, set a high performance bar, but its high cost limited widespread adoption.

The move to release Sol and Luna at significantly reduced prices follows trends in the industry towards democratizing AI access, similar to recent offerings by competitors like Anthropic. The focus on cost reduction through improved caching and inference aligns with broader efforts to improve efficiency without sacrificing too much quality.

Independent evaluations, such as Artificial Analysis, have tracked these models’ performance and cost metrics, noting that while the new models maintain a competitive edge in some benchmarks, they also exhibit regressions in others, especially in knowledge and presentation quality.

Amazon

AI development GPU server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Performance

It remains unclear how the reduced presentation quality will affect long-term use cases requiring detailed, accurate outputs. The impact of increased refusal rates on workflows that depend on comprehensive responses needs further investigation. Additionally, the stability of benchmark scores over time and across different tasks is still being monitored, with some evaluations indicating regressions.

Further testing is needed to determine whether these models’ performance variations are temporary or indicative of broader limitations introduced by the new cost-saving techniques.

Amazon

AI inference optimization tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in Model Evaluation and Adoption

Industry stakeholders are expected to conduct extensive testing of GPT‑6 Sol and Luna across diverse applications, focusing on accuracy, hallucination rates, and presentation quality. OpenAI is likely to update its documentation and provide more detailed benchmarking data as real-world deployment progresses.

Monitoring of benchmark stability and performance regressions will continue, with potential model refinements or tuning to address identified weaknesses. Meanwhile, organizations will weigh cost savings against potential quality trade-offs before integrating these models into critical workflows.

Amazon

cost-effective AI computing hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How much cheaper are GPT‑6 Sol and Luna compared to previous models?

GPT‑6 Sol costs about half as much per 1 million tokens as GPT‑5.6, with input tokens at $2.00 and output tokens at $10.00. Luna is even more affordable, at $0.10 and $0.50 respectively, representing roughly a 50-60% cost reduction.

Do the new models perform worse or better than earlier versions?

Cost-wise, they are more economical, but some evaluations show regressions in knowledge work and presentation quality, especially in detailed or complex outputs. Hallucination rates have decreased, but at the expense of more refusals, which may impact certain workflows.

What are the main trade-offs with these models?

The models reduce hallucinations and costs but may produce less detailed or complete outputs, especially in knowledge-intensive tasks. Increased refusal rates can also affect workflows that rely on comprehensive responses.

Will these models replace high-end models like Astra?

Not necessarily. Astra remains the top of the range for tasks requiring the highest accuracy and detail. Sol and Luna are positioned as more accessible, cost-effective options suitable for broader deployment where some quality trade-offs are acceptable.

How might these models influence AI adoption industry-wide?

The significant cost reductions could lower barriers for many organizations to incorporate AI into their operations, potentially accelerating automation and digital transformation across sectors.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Technology operations signal monitor: Show HN: Kage – Shadow any website to a single binary for offline viewing

Kage, a new tool allowing users to shadow websites into a single binary for offline viewing, is gaining interest among product and engineering leads at small software firms.

RHEO on Steam: One Toy, Every Screen

RHEO is launching on Steam, offering a seamless, cross-device fluid art experience on PC, Steam Deck, VR, and more with one purchase.

Waves, Not a Wall: Inside DeepMind’s Map From AGI to Superintelligence

DeepMind researchers publish a detailed framework outlining pathways from human-level AI to superintelligence, highlighting potential progress routes and challenges.

Staking Mechanisms in Ethereum ETFS: Technical Integration With Staking Rewards

Discover how staking mechanisms in Ethereum ETFs integrate technical features that reward investors and support the network’s security—continue reading to learn more.