Grok 4.7, Opus 5.5, GPT-6 Sol and Luna: Four Models in Two Days
TL;DR
xAI released Grok 4.7 on September 21, 2026 at an unchanged $2/$6 per million tokens. On September 22 Anthropic shipped Claude Opus 5.5 at $4/$20, 20% below Opus 5, and about 90 minutes later OpenAI released GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50). xAI's headline table skips GPT-6 Astra, some OpenAI charts use older or lower-effort Anthropic models, and Anthropic's announcement leaves out results Astra wins. Artificial Analysis found that heavy token use cancels much of the per-token savings for Grok 4.7 and Opus 5.5 at their highest effort settings.
Three labs shipped four models in two days. xAI released Grok 4.7 on Monday, September 21. On Tuesday, Anthropic launched Claude Opus 5.5, and about 90 minutes later OpenAI followed with GPT-6 Sol and GPT-6 Luna. Nobody raised prices, and every launch’s benchmark charts were out of date within a day. Here’s what each model costs, what it breaks, and which numbers are vendor claims.
xAI’s Grok 4.7 is late, and cheaper per token than per task
xAI, which now also trades as SpaceXAI, calls Grok 4.7 its “most capable model for coding and knowledge work,” built on a new, larger base model than Grok 4.6. It’s late: Decrypt counted at least five walk-backs of Elon Musk’s timeline since late July, the last a “needs a few more days to cook” on September 11.
Pricing is unchanged from Grok 4.6. It’s $2 input, $0.50 cached and $6 output per million tokens below 200K prompt tokens, and $4/$1/$12 on the whole request once a prompt crosses that line. There’s no Batch API discount. The model page lists a 500K-token context window, text and image input, reasoning effort from low to xhigh, and both the Responses and Chat Completions APIs.
It’s grok-4.7 on the API, the default in xAI’s Grok Build agent and Office add-ins, in Cursor, and rolling out in GitHub Copilot. The Grok apps and X get it “at a later date,” says the model card. The announcement doesn’t link the card, which reports regressions such as self-harm compliance rising from 0.84% to 1.05%.
The benchmarks are xAI’s claims. The headline table pits Grok 4.7 against Claude Fable 5.1 and OpenAI’s GPT-5.6 Sol. GPT-6 Astra, OpenAI’s flagship since September 3, appears only in a secondary chart. Even in the headline table, Grok 4.7 trails Fable 5.1 on CursorBench and Terminal-Bench. CursorBench comes from Cursor, which SpaceXAI now owns and whose workflow data helped train Grok 4.7. Artificial Analysis counted about 81,000 output tokens per index task, against Astra’s 27,000. It puts the cost per task at $3.74 against Astra’s $3.26, even though Grok’s output tokens cost an eighth as much.
Anthropic’s Claude Opus 5.5 brings Fable-level claims to a cheaper tier
Anthropic says Opus 5.5, the first model in its Claude 5.5 family, “performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.” Fable 5.1 stays at the top of the price list, and the models overview still recommends it for demanding reasoning and long-horizon agentic work.
Pricing is $4 input and $20 output per million tokens, 20% below Opus 5 and 60% below Fable 5.1. Cache reads cost $0.20, only 20% less than on Fable 5.1, and there’s no surcharge across the 1M-token window. Output tops out at 128K tokens, the knowledge cutoff is June 2026, and adaptive thinking is always on. Anthropic says the “40% less” applies at default settings, where Opus 5.5 runs at medium effort and Opus 5 ran at high. At max effort, Artificial Analysis counted about 119,000 output tokens per index task, against 73,000 for Opus 5 and 27,000 for Astra. That left Opus 5.5’s cost per task level with Opus 5’s.
It’s claude-opus-5-5 on the Claude API, the major clouds, GitHub Copilot (Pro+ and up) and paid claude.ai plans, and the default in Claude Code v2.1.280. On top of what Opus 5 already rejected, the migration guide lists four breaking changes. Thinking can’t be disabled, forced tool_choice is rejected, and the Claude API refuses the older computer_20251124 tool. Thinking blocks are also bound to the model and conversation, as on Fable 5.1.
It’s Anthropic’s first release since it called for pacing the frontier. METR and other external evaluators tested it before launch, and the system card says US CAISI did too. The card says that without product mitigations, the model acted on malicious instructions planted in pasted text in about 2% of attempts, which Opus 5 never did. Anthropic also sees signs that the model “often suspects it is being evaluated.”
The benchmark table does include OpenAI’s flagship, GPT-6 Astra, which wins two of its rows. Its CursorBench cost claim, though, is made against the older GPT-5.6 Sol, because Astra has no score there. The system card shows Astra ahead on FrontierSWE v2 (65.5% to 62.3%), which the announcement leaves out. On Terminal-Bench 4.0, Artificial Analysis measured 59.6%, level with Astra, against Anthropic’s own 66.4%.
OpenAI’s GPT-6 Sol and Luna make GPT-6 cheaper to run
OpenAI’s GPT-6 guide splits the family three ways: Astra for maximum capability, Sol for strong reasoning on demanding tasks, and Luna for efficient, repeatable work at scale.
Pricing for Sol is $2 input, $0.20 cached and $10 output per million tokens. Luna is $0.10, $0.01 and $0.50. Once a prompt passes 272K tokens, the whole request is billed at higher rates. OpenAI puts that at 50% below GPT-5.6’s promotional prices, which had already cut GPT-5.6 Sol to $4/$20. VentureBeat reports that an OpenAI spokesperson called the new prices permanent. Both models have a 1.05M-token context window, 128K output, image input, and reasoning effort from none to max.
They’re gpt-6-sol and gpt-6-luna in the API, and also on Microsoft Foundry, Amazon Bedrock and GitHub Copilot, where Copilot Pro gets Luna but not Sol. In ChatGPT they live in Work and Codex, not regular chat. Free and Go users get only Luna, and only in the desktop app.
The API rules carry over from GPT-5.6, so code written for older models will hit them. In Chat Completions, function calling works only with reasoning effort set to none, so tools with reasoning mean the Responses API. With reasoning on, you have to drop temperature, top_p and logprobs, and there’s no minimal effort. Safety documentation is lighter than Astra’s. There’s no standalone system card, just an appendix to the GPT-6 Astra card. It rates both models High for cybersecurity and biology under OpenAI’s Preparedness Framework, and it names no external evaluators.
OpenAI’s rival comparisons are all against Anthropic, and some use older or lower-effort models. According to VentureBeat, OSWorld 2.0 puts Sol at xhigh against Opus 5 at medium effort, and Sol roughly ties it, 60.5% to 60.3%, at about 80% lower cost per task. DeepSWE compares against Fable 5 rather than Fable 5.1, and Sol loses, 68.8% to 69.9%, again at about 80% lower cost. Other charts do use Fable 5.1, but none includes Opus 5.5.
What this means
Per-token prices fell or held everywhere, but the Grok 4.7 and Opus 5.5 numbers above show how a token-hungry model hands the savings back per task. On Artificial Analysis’s independent Intelligence Index, Opus 5.5 scores 58, GPT-6 Astra and Fable 5.1 score 53, Sol 48, Grok 4.7 46 and Luna 37.
Release timing works as competitive strategy whether or not anyone planned it. With launches a day or 90 minutes apart, no chart includes the rival that shipped next. Because Anthropic went first, OpenAI’s day-one charts compared Sol with Opus 5, which Anthropic had just replaced, and the two launches shared headlines.
If you’re choosing, run your own tasks at the effort level you’ll ship and measure the cost of each completed task. Read the migration guide before swapping a model ID, because both Opus 5.5 and GPT-6 reject requests that older models accepted. For high-volume extraction and summarization, Luna costs a twentieth of Sol per token, far below anything else here. For agentic coding, test Sol and Opus 5.5 against each other on your own repo. Below 200K prompt tokens, Grok 4.7 has the cheapest output of the three coding models, so check how many tokens it uses before you count the savings.
FAQ
When were Grok 4.7, Claude Opus 5.5 and GPT-6 Sol released?
xAI released Grok 4.7 on Monday, September 21, 2026. Anthropic released Claude Opus 5.5 on Tuesday, September 22, 2026, and OpenAI released GPT-6 Sol and GPT-6 Luna about 90 minutes later the same day, at 11:00 a.m. Pacific time.
How much do Grok 4.7, Claude Opus 5.5 and GPT-6 Sol cost?
Per million tokens: Grok 4.7 costs $2 input and $6 output below 200K prompt tokens, doubling above that. Claude Opus 5.5 costs $4 input and $20 output, with cache reads at $0.20. GPT-6 Sol costs $2 input and $10 output, and GPT-6 Luna costs $0.10 input and $0.50 output, with higher rates for prompts above 272K tokens.
What is the difference between GPT-6 Sol and GPT-6 Luna?
OpenAI positions GPT-6 Sol for strong reasoning on demanding tasks, including complex coding and agentic workflows, and GPT-6 Luna for efficient, repeatable work at scale, such as summarization and extraction. Both have a 1.05M-token context window and 128K output tokens. Luna is 20 times cheaper per token. GPT-6 Astra remains OpenAI's most capable model.
Is Claude Opus 5.5 better than Claude Fable 5.1?
Anthropic says Opus 5.5 performs at the level of Fable 5.1 on most work. Its input and output prices are 60% below Fable 5.1's, though cache reads are only 20% cheaper, and it scores higher on Artificial Analysis's Intelligence Index (58 versus 53). Anthropic still recommends Fable 5.1 for demanding reasoning and long-horizon agentic work, and says the real-world gap is narrower than benchmark scores suggest.