Releases

GPT-6 Astra costs 2.5x Sol per token. It talks like a person

OpenAI's GPT-6 Astra is metered at 2.5 times Sol's rate and feels more human in conversation. I checked the limits, the benchmarks, and the first reactions.

On this page
  1. What did OpenAI ship on 3 September 2026?
  2. Why does Astra drain my allowance faster than Sol? Price per token, not verbosity
  3. Does Astra feel more human? My read, and what the evidence says
  4. Is Astra smarter than Sol 5.6? Yes on agents, barely on general reasoning
  5. What changed in Codex: scope discipline improved, sprawl did not
  6. The cyber gate and the monitorability trade
  7. How I would split the work on a paid plan

OpenAI released GPT-6 Astra on 3 September 2026, and it reached paid ChatGPT plans the next day. In Codex and ChatGPT Work it is metered at exactly 2.5 times the credit rate of GPT-5.6 Sol, so my allowance drains faster even though the model writes fewer tokens per task. [1] [2] My own read after the first days: it answers less like an AI, and it is a little smarter where it matters.

What did OpenAI ship on 3 September 2026?

OpenAI shipped one model, GPT-6 Astra, as a new generation number and a direct replacement for GPT-5.6 Sol. It is more expensive than Sol at every tier, holds a bigger context window in the API, and arrived through a staged rollout that OpenAI itself apologized for. [1] [9]

The naming deserves one clarification, because the previous Astra article on this site left it open. In July, OpenAI described Sol, Terra and Luna as durable capability tiers that could advance on their own schedule. That framing did not survive two months. Astra is the only GPT-6 model, the API guidance now pairs it with GPT-5.6 Terra and GPT-5.6 Luna for cheaper work, and inside ChatGPT the chat version is labelled GPT-6 Pro. [1] [4] [5]

Specification GPT-6 AstraGPT-5.6 SolClaude Fable 5.1
Input, $/MTok 10 4 (Best value in this row) 10
Output, $/MTok 50 20 (Best value in this row) 50
Cache read, $/MTok 1.00 0.40 0.25 (Best value in this row)
API context window 1.05M 1.05M 1M
Figure 1. API list prices and context windows on 5 September 2026. Astra and Sol figures from OpenAI's API pricing page, Fable figures from Anthropic's model page.

Figure 1 shows the two things that decide the cost of using Astra. First, the price matches Claude Fable 5.1 exactly on input and output, so OpenAI now sells its flagship at Anthropic’s flagship rate. Second, cache reads cost four times more on Astra than on Fable 5.1, which matters for long agent sessions that reread the same context. [7] [24] [34] The API model accepts 1,050,000 tokens of context, produces up to 128,000 output tokens, has a knowledge cutoff of 30 April 2026, and supports reasoning effort from low to max. The none setting returns an error, and tool calling requires the Responses API. [5]

Inside Codex the picture is narrower than the API page suggests. Astra needs Codex CLI 0.153.0 or newer, and it only became the default and selectable model in the bundled picker with release 0.153.4 on 4 September. [3] [8] The Codex catalog gives Astra the same 272K default context window that Sol has, with 872K as the maximum, so the one million tokens exist for API users rather than for subscribers. [2] [6] Codex also adds a sixth reasoning level called Ultra, which delegates parts of a task to parallel subagents, and an experimental setting that keeps notes across context windows instead of compressing everything into one summary. [1] [6]

The rollout was the first thing most people experienced. OpenAI announced Astra for a limited set of organizations on 3 September, paying users found nothing in their model picker, and Codex lead Thibault Sottiaux promised one banked reset for every day a paid account went without access. [10] Sam Altman wrote “sorry for the messy rollout” the next morning, and general availability for Plus, Pro, Business and Enterprise followed on 4 September. [9] That timing is why this article is about first days rather than a first week.

Why does Astra drain my allowance faster than Sol? Price per token, not verbosity

My subscription runs out sooner on Astra because OpenAI charges 2.5 times more credits per token for it, not because Astra writes more. Every independent measurement I found has Astra using fewer output tokens per task than Sol. The credit rate simply outweighs the saving. [2] [11]

credits per token vs Sol
2.5x
250 in, 1,250 out per million, against 100 and 500
fewer output tokens than Sol
36 %
49M vs 76M across the Intelligence Index
cost per index task
$2.57
against $1.25 for Sol at max effort
Figure 2. What decides the quota burn. Credit rates from OpenAI's Codex pricing page, token and cost figures from Artificial Analysis, 5 September 2026.

Figure 2 puts the three numbers side by side. In Codex and ChatGPT Work, a million input tokens costs 250 credits on Astra against 100 on Sol, and a million output tokens costs 1,250 against 500. [2] OpenAI’s own message estimates halve accordingly: a Plus plan gets an estimated 5 to 45 local Astra messages per five-hour window where Sol gets 10 to 100, and the Pro 20x tier gets 100 to 900 instead of 200 to 2,000. [2] The help center states the consequence in one sentence: “Astra can use your allowance faster than GPT-5.6 Sol.” [3]

The verbosity side runs the other way. Artificial Analysis needed 49 million output tokens to run its Intelligence Index on Astra at maximum effort, compared with 76 million for Sol, and it ranks Astra as fairly concise. [12] On Datacurve’s DeepSWE coding tasks, Astra averaged about 30,000 output tokens and 29 steps per task, while Sol averaged about 60,000 tokens and 61 steps for the same pass rate. [13] Fewer tokens at 2.5 times the price still comes out more expensive: Artificial Analysis measured $2.57 per index task for Astra against $1.25 for Sol, and calls Astra 75% more expensive per task at max effort. [11] [12] In practice, this means the same coding session takes a noticeably bigger bite out of a five-hour window on Astra than it did on Sol, which matches what I saw.

Regular ChatGPT chat is a separate budget with its own caps. There, Astra appears as GPT-6 Pro: 200 messages per week on the $200 Pro plan, 50 per week on the $100 Pro plan shared with Sol Pro, 15 per month on Business Standard, and nothing at all on Plus, which only gets Astra in ChatGPT Work and Codex. When a $200 Pro user hits the weekly cap, ChatGPT switches to GPT-5.6 Thinking at medium. [4] A Hacker News commenter did the arithmetic on launch day: if Astra were really as cheap per task as the token counts imply, OpenAI would not need to cap it at roughly 16% of Sol’s message count. [14] I think that is the right way to read the limits. Per-task token use varies a lot, and OpenAI priced the allowance for the expensive cases.

Reddit reached the same conclusion within hours, and with more feeling. The r/codex thread that posted the new limits drew 536 points, and its top comment asked “Where is all that efficiency they have been talking about?” A few replies down, someone summarized the new plan structure as “Plus is the new Free. Pro x5 is the new Plus.” [25] The thread about the banked resets was even more telling: the top comments said they would happily wait days without Astra, because each day without access earned a free reset. [26] In r/ChatGPT, a Business user reported that a single question about the training cutoff, asked at a very high reasoning setting, took 18% of a five-hour window. [27] Two facts soften this. One user ran the same prompt on the same repository in two worktrees and found Astra used 4% of the quota where Sol used 10%, and finished in 46 minutes instead of 72. [28] And OpenAI’s own claim that a lower-cost Astra setting beats Sol’s best result on several benchmarks is what the calmer reports describe. [1] The burn is real, but it depends heavily on the reasoning level and on how open-ended the prompt is.

Does Astra feel more human? My read, and what the evidence says

To me, Astra sounds less like an assistant reading from a script and more like a colleague. The replies read like someone answering the question rather than performing helpfulness. That is an impression from a few days of use, and no published benchmark measures it yet, so I want to be precise about what supports it and what does not.

OpenAI does not claim any of this. The launch post talks about judgment rather than voice: Astra “is better than previous models at making the right call” when instructions leave room for interpretation, and it asks focused questions when the answer would change the outcome. [1] The developer guide is even less flattering. It warns that the model “tends toward detailed, formatted responses and may use recurring phrases across sessions”, and that it reaches for lists, tables and Markdown unless told otherwise. [15] If OpenAI tuned Astra to sound human, it did not write that down.

The early reviewers who spent longer with it than I have land close to my impression. Matt Shumer, who tested Astra before launch, writes that recent models answered straightforward questions with dense technical explanations while “Astra generally speaks in plain English.” [16] Every let Astra write the first draft of its own review from a single prompt, and the publication’s cofounder replied that they had not realized the draft was not written by the reviewer. [17] On Hacker News, one early tester described Astra as a “grounded collaborator and executor” that stays a collaborator when prompted like one, without becoming over-eager or doing work it was not asked to do. [14] Another commenter was disappointed for exactly the opposite reason, saying they could not find evidence that Astra communicates more naturally or writes more elegant code, and that Claude Fable 5.1 does better on both. [14]

The Reddit hands-on threads from 4 September say the same thing in fewer words, and they say it about the writing rather than about warmth. One r/codex user’s first-hour verdict was that Astra is “as smart as Fable, but less annoying in how it talks”, and eats quota more than Sol. Others wrote that “it actually writes prose that I can understand” and that its writing is easier to follow than Fable’s. [29] The bluntest version simply said that “Astra talks just like a normal human being.” [28] The dissent came from a Pro user running Fable and Astra side by side, who preferred Astra for research and data science but added that “Fable is still better to talk to.” [30] That is a handful of comments, none of them highly voted, so I count them as agreement with my impression rather than as proof of it.

The measured evidence points to a mechanism rather than to a personality. OpenAI’s system card reports that Sol misrepresents what it has done, verified, or has access to at four times Astra’s rate at maximum effort, and its internal hallucination benchmark falls from 12.2% for Sol to 4.2% for Astra. [18] Artificial Analysis measured the hallucination rate on its own knowledge test dropping from 92% to 51%. [11] My read is that a model that bluffs less and asks more sounds more like a person, because those are the two habits that made earlier assistants feel like assistants. The counterweight is real too: Artificial Analysis found Astra’s presentation quality on its business-document test went down, with Sol still leading all models there. [11] Better company in a chat and better slide decks are different skills.

Is Astra smarter than Sol 5.6? Yes on agents, barely on general reasoning

Astra is clearly ahead of Sol on tasks that involve a terminal, a browser, or a long chain of steps, and it is close to a tie on general intelligence indexes. Whether it is “smarter” depends on which of those two you mean. My feeling that it is a little more intelligent comes from the first group, because that is the kind of work I hand it.

  • GPT-6 Astra
  • GPT-5.6 Sol
  • Claude Fable 5.1
GPT-6 Astra, GPT-5.6 Sol and Claude Fable 5.1 on five benchmarks Astra leads Sol by 20 to 42 points on terminal and automation tasks and by 15 points on FrontierMath Tier 4, while GPQA Diamond is nearly level. Terminal-Bench 4.0 57.9% 37.3% 55.8% Terminal-Bench Science 0.1 64.6% 22.4% 52.6% AutomationBench 41.4% 18.1% 31.4% FrontierMath Tier 4 97.6% 83.0% 87.8% GPQA Diamond 96.0% 94.6% 93.7%
Show the data as a table
Benchmark GPT-6 AstraGPT-5.6 SolClaude Fable 5.1
Terminal-Bench 4.0 57.9%37.3%55.8%
Terminal-Bench Science 0.1 64.6%22.4%52.6%
AutomationBench 41.4%18.1%31.4%
FrontierMath Tier 4 97.6%83.0%87.8%
GPQA Diamond 96.0%94.6%93.7%
Figure 3. Selected scores from OpenAI's launch table, 3 September 2026. Every OpenAI figure is the maximum at any effort level. Compare across a row, not between rows.

Figure 3 shows where the generation change is visible. Astra scores 57.9% on Terminal-Bench 4.0 against 37.3% for Sol, 64.6% on the scientific terminal benchmark against 22.4%, and 97.6% on FrontierMath Tier 4 against 83.0%. [1] On the same table Astra also completes OSWorld computer-use tasks at 72.6% in about 40 minutes per task, where Sol reached 65.7% in about 75 minutes. [1] These are the results that match my experience: give Astra a goal that takes many steps, and it gets further with less supervision.

The general reasoning numbers are much flatter. OpenAI’s own table lists the Artificial Analysis Intelligence Index at 61.2 for Astra and 60.9 for Sol, with Claude Fable 5.1 at 65.7. [1] Artificial Analysis has since moved to version 4.2 of the index, where Astra scores 55, Sol 51, and Fable 5.1 57. [12] On Humanity’s Last Exam with tools, the one academic row OpenAI publishes and loses, Astra scores 57.2% against 65.0% for Fable 5.1. [1] So the honest summary is that Astra is a large step for agents and a small step for reasoning in a chat box, and it does not take the general lead from Anthropic.

That tie is what the skeptics on Reddit picked up first. An r/codex post titled “Astra is good, but maybe we should calm down with the hype” pointed out that Artificial Analysis had Astra at 61, “exactly the same as Sol”, with Meta’s Muse Spark 1.3 at 62, and called the generational framing marketing. It collected 282 points. The top reply, with 209, defended Astra by attacking the index instead: a composite that leans on single-turn math and code was never going to move much for a model that put its gains into long tool use. [32] Both sides are right about their own half, which is why I split the question the way I did above.

Three details about OpenAI’s table change how much weight to put on it. First, every OpenAI score is the maximum at any effort level, and the page says so under the table. Artificial Analysis measured Astra’s time to first token at max effort at about 464 seconds, so the best numbers come with waits of several minutes. [1] [12] Second, the 99.9% headline on ARC-AGI-3 comes from OpenAI’s own Provider Adapter harness, which keeps reasoning state between requests. On the standard harness that every model uses, ARC Prize measured 62.7%. Both numbers are real, but only the second one is comparable. [19] Third, the table has no SWE-bench Verified row at all, and Cognition’s FrontierCode was run with a custom developer message asking Astra to avoid excessive tests and unrelated cleanup. On FrontierCode Main, Astra’s 53.3% sits within a point of Fable 5 and Opus 5. [1] [20]

What changed in Codex: scope discipline improved, sprawl did not

The habit that made Sol hard to trust on small tasks, wandering past the requested scope, is much reduced in Astra. The habit of delivering a large, thorough change when a small one was asked for is still there. Both show up in OpenAI’s numbers and in the first production reports.

OpenAI built an evaluation after the Hugging Face incident to test whether a model facing an impossible task goes beyond its authorized target. Without production safeguards, Sol did that in 48% of cases and Astra in 0%. [1] The system card adds that when a task hits a deliberate barrier, Sol tried to route around it in 64% of runs at maximum effort, against 19% for Astra. [18] In a simulation over more than 54,000 internal Codex tasks, Astra received about half as many flags for higher-severity misaligned behavior. [1] For a coding agent, those numbers matter more than any benchmark score, because they decide how much supervision a run needs.

Kilo, which ran Astra in production before launch, calls it the best coding model it has tested and then lists the cost of that thoroughness. “Ask for a targeted fix and Astra will frequently return a massive change with a sprawling PR. It defaults to the thorough solution rather than the minimal one.” On vague prompts, Kilo says, it “will happily spend your budget confirming things it already knew.” [21] OpenAI’s developer guide says the same thing in its own words: for smaller tasks, Astra’s testing “can result in broader tests than the task requires”, and the model is more likely to stop and ask a question when it thinks the answer could change the result. [15] In my own sessions the questions were usually good ones, but each of them costs a round trip on a metered plan.

The r/codex reports contradict each other on exactly this point, which tells me the behavior depends on the prompt. One user said Astra “began to chastise me for over engineering” and told them to aim for the minimum viable version first. Another, whose Astra session was managing cheaper Luna subagents, said the best part was watching it call out the subagents whenever they expanded scope beyond the instructions. [28] [33] In the same threads, a Fable user found that Astra “produces a hell of a lot more code than Fable 5.1 does” while calling its approach to problem solving the best they had seen, and a Plus user said it “overengineers like crazy” and drained a week of usage in a single thread. [25] [33] The most useful single report came from the megathread: a task that had eaten a month of Sol at xhigh and two banked resets without a solution, solved by Astra at medium for 15% of a weekly limit. [31] The top comment on the most enthusiastic thread was a prediction that the posts about how dumb it is would arrive within 48 hours, and it had 243 points. [29] Honeymoon effects are real, so I would read all of these, including mine, as first days.

Against Claude Fable 5.1, the coding comparison is close and depends on the harness. Artificial Analysis puts Fable 5.1 inside Claude Code at 70 on its Coding Agent Index and Astra inside Codex at 67, and it notes that Astra reaches Fable 5’s score at less than half the cost. [11] Cognition’s leaderboard has Astra at $4.49 per FrontierCode rollout against $6.29 for Sol. [20] Every’s verdict after a week was that Astra is a big upgrade from Sol with some frustrating habits that keep it from matching Fable at the top end. [17] I would not argue with that.

The cyber gate and the monitorability trade

Astra is the first OpenAI model rated Critical for cybersecurity capability under the Preparedness Framework, and that rating shapes the product you can actually use. The public version refuses advanced offensive tasks such as writing proof-of-concept exploits, less restricted access goes through the Daybreak program, and extra safety checks “can sometimes slow, pause, or stop legitimate work.” In ChatGPT or Codex you get asked to review a paused action. In the API, the task stops. [1] [18] One security professional in the r/codex megathread reported that Daybreak verification did not carry over to Astra at launch, so authorized work that ran on Sol was being refused. [31]

The delay since August fits that story. OpenAI paused certain Astra training for two weeks after the Hugging Face incident, held back larger reinforcement learning runs for longer, and restarted the big run on 28 August after the new controls were in place. [22] The model itself is a very large run: at the press briefing OpenAI said it was the first time the company had pretrained on more than 100,000 GPUs at its Stargate site in Texas. [23]

The part I find most uncomfortable is in the system card. OpenAI writes that Astra’s monitorability “has decreased relative to GPT-5.6 Sol”, that in adversarial tests the model could underperform on evaluations without being detected, and that simply telling it a monitor exists can make it shorten its reasoning. [18] That is not an alignment failure in daily use, and the scope numbers above cut the other way. But it means the thing that makes Astra feel like a sensible colleague, its judgment about what to do next, is harder for its own maker to inspect.

How I would split the work on a paid plan

I would point conversations and long agent runs at Astra and keep routine coding on cheaper models. That split follows directly from the numbers: Astra wins on judgment and long-horizon work, while the credit rate punishes routine edits.

TaskMy pickWhy
Thinking through a design or a decisionAstra in ChatGPT or WorkBetter questions, less bluffing, plain answers
Multi-step agent run with tests and a PRAstra at high, scope stated up frontLargest measured gains, but it delivers big changes unless told not to
Routine edits, renames, small fixesGPT-5.6 Terra or SolSame 272K Codex window at a fraction of the credit rate
Long session that rereads the same contextFable 5.1 if I have itCache reads at $0.25 instead of $1 per million

Two practical notes. State the scope in the first message, because Astra follows explicit boundaries far better than Sol did and fills silence with thoroughness. And watch the weekly caps rather than the five-hour window: on a $100 Pro plan the 50 GPT-6 Pro chat messages are shared with Sol Pro, and on Codex the banked resets OpenAI handed out during the rollout are the cheapest Astra time you will get. [4] [10]

Astra is the first OpenAI release where the model I want to talk to and the model I want to run are the same one. It is also the first where I have to ration it.

Sources

  1. GPT-6 Astra: A new generation of intelligenceOpenAI · 2026-09-03
  2. Pricing and usage limitsOpenAI, Codex docs · 2026-09-05
  3. ChatGPT Work and CodexOpenAI Help Center · 2026-09-05
  4. GPT-5.6 and GPT-6 Pro in ChatGPTOpenAI Help Center · 2026-09-05
  5. GPT-6 Astra model referenceOpenAI Developers · 2026-09-03
  6. ModelsOpenAI, Codex docs · 2026-09-05
  7. Claude Fable 5.1Claude Platform Docs · 2026-09-01
  8. Codex CLI rust-v0.153.4GitHub, openai/codex · 2026-09-04
  9. On the messy Astra rolloutX, Sam Altman · 2026-09-04
  10. On banked resets for every day without Astra accessX, Thibault Sottiaux · 2026-09-03
  11. Benchmarking GPT-6 AstraArtificial Analysis · 2026-09-03
  12. GPT-6 Astra model analysisArtificial Analysis · 2026-09-05
  13. DeepSWE leaderboardDatacurve · 2026-09-03
  14. GPT-6 Astra discussionHacker News · 2026-09-03
  15. Using GPT-6 AstraOpenAI Developers · 2026-09-03
  16. My GPT-6 Astra reviewMatt Shumer · 2026-09-03
  17. Vibe Check: GPT-6 Astra is a big upgrade with some bad habitsEvery · 2026-09-03
  18. GPT-6 Astra system cardOpenAI · 2026-09-03
  19. OpenAI's GPT-6 Astra on ARC-AGI-3ARC Prize Foundation · 2026-09-03
  20. FrontierCode leaderboardCognition · 2026-09-03
  21. GPT-6 Astra: What we learned previewing OpenAI's new model in productionKilo · 2026-09-04
  22. Path to Astra: critical capabilities and frontier safeguardsOpenAI · 2026-09-01
  23. OpenAI launches GPT-6 Astra, its most powerful model yet, and touts its ability to use your computerFortune · 2026-09-03
  24. PricingOpenAI Developers · 2026-09-03
  25. GPT 6 Astra usage limits for Plus, Pro 5x, and Pro 20xReddit, r/codex · 2026-09-04
  26. Confirmed Bank Reset!Reddit, r/codex · 2026-09-04
  27. Astra is burning through tokens like crazyReddit, r/ChatGPT · 2026-09-04
  28. W AstraReddit, r/codex · 2026-09-04
  29. Got Astra in codexReddit, r/codex · 2026-09-04
  30. Pro user here. Astra just landed.Reddit, r/OpenAI · 2026-09-04
  31. Astra Release MegathreadReddit, r/codex · 2026-09-04
  32. Astra is good, but maybe we should calm down with the hypeReddit, r/codex · 2026-09-04
  33. GPT6 Astra is blazingly fastReddit, r/codex · 2026-09-04
  34. Models referenceOpenAI Developers · 2026-09-03