Launched in early September 2026, GPT-6 Astra shows that OpenAI is no longer competing solely on the ability to answer questions or write code. The new focus is on AI that can operate a computer by itself, use tools and complete a workflow from start to finish.

On 3 September 2026, OpenAI officially introduced GPT-6 Astra, a new model focused heavily on coding, research, computer use and multi-step tasks. If earlier GPT generations were often judged by their ability to answer questions or generate code, Astra points to a clearer direction of development: AI Agents that can work directly rather than merely instructing users.
This is also what makes Astra a notable rival to Claude Fable 5.1, Anthropic's new model, which is already very strong at coding, knowledge work and long-running tasks.
Astra's biggest difference does not lie in a single reasoning benchmark, but in its ability to combine multiple tools within the same workflow.
Astra is designed to work with browsers, terminals, documents, spreadsheets and other software. Instead of users asking the AI and then carrying out each step themselves, the goal is to hand the AI a task and let the agent handle most of the process on its own.
For example, a developer could ask it to:
The important point here is not how elegant a piece of code the AI can write, but how many steps it can complete on its own before a human needs to intervene.
According to figures published by OpenAI, GPT-6 Astra scores 41.4% on AutomationBench, while Claude Fable 5.1 scores 31.4%. This is a benchmark focused on real work workflows, and the gap suggests Astra holds a significant advantage in agentic workflows.
Looking at coding alone, the race remains far more balanced.
Claude Fable 5.1 remains a very strong model for tasks such as analysing large codebases, debugging, root-cause analysis and long-running coding. Anthropic reports that Fable 5.1 scores 55.8% on Terminal-Bench 4.0 and 73.4% on CursorBench 3.2.
OpenAI, meanwhile, positions Astra as strong in software engineering, especially when coding is combined with terminals, browsers and other tools.
This creates a rather interesting distinction.
Claude tends to stand out when deep analysis is needed and a line of reasoning must be maintained across a long task. Astra, in turn, is particularly notable when a task requires the AI both to reason and to operate multiple tools directly in order to complete the work.
For developers, this may matter more than comparing a few benchmark points.
A model that writes code 5% better but has to keep asking the user again may not be as effective as an agent that can build, test, read logs and fix bugs across many consecutive rounds on its own.
GPT-6 Astra does not win on every front.
Even in the figures published by OpenAI, Claude Fable 5.1 scores 65.7 on the Artificial Analysis Intelligence Index, higher than Astra's 61.2.
Anthropic is also focusing heavily on scientific research, knowledge work and tasks that run for many hours. Fable 5.1 is designed to maintain a working state for longer, check its own results and handle complex problems without easily losing the thread.
So concluding that GPT-6 Astra has "beaten Claude" would be premature.
A more reasonable reading is that the two companies are emphasising two slightly different strengths:
The AI race used to revolve around questions such as: “Which model is smarter?”
“Which model writes code better?”
“Which model has a larger context?”
But with Astra, the criteria for evaluation may be shifting.
The more practical question will be:
Which model can complete more work on its own with less human intervention?
This is a huge change.
In enterprise environments, the cost of AI is not just the price of tokens. The cost also includes the human time spent checking, rewriting prompts, re-running, explaining context and dealing with an agent that goes off track.
If a model costs a little more but can complete a task after a single handover, it may well be cheaper than a model that needs five or ten rounds of back-and-forth.
At present, there is no outright winner.
If the work is mainly document analysis, long-form reasoning, research, or requires a coding partner that holds context well over a long period, Claude Fable 5.1 remains a very strong choice.
If the work requires combining code, a terminal, a browser and other applications to carry out a complete workflow, GPT-6 Astra is a particularly noteworthy model.
For developers or teams building AI Agents, Astra may be more appealing precisely because of its ability to turn reasoning into action.
GPT-6 Astra is not just an upgrade to GPT's “intelligence”.
The most notable point is that OpenAI is pushing hard on the shift from chatbot to AI Agent capable of directly carrying out work.
Claude remains an extremely strong competitor and even leads Astra on some benchmarks. But Astra shows that the AI race is entering a different phase: the best model may not be the one that answers best, but the one that completes real work best.
In the next few years, this will probably be a far more important question than how many benchmark points GPT or Claude has gained.
Comments
No comments yet. Be the first!
You need to sign in to comment.