OpenAI released GPT-6 Astra on September 3 at $10 per million input tokens and $50 per million output, 2.5 times the price of the model it replaces.
Testers with early access posted a street-by-street Manhattan in Unreal Engine, a browser-based 3D Hangzhou built in 24 minutes, a multiplayer shooter made in a day, and a Bach chorale with no voice-leading errors.
The same testers rated its writing below its own predecessor, and Artificial Analysis measured a drop of roughly 80 Elo points on a benchmark of economically valuable professional work.
OpenAI released GPT-6 Astra on September 3, and within 48 hours the developers who got early access had turned the launch into a public stress test. What they posted splits along one clean line.
Astra is the strongest model anyone has used for anything spatial, mechanical, or agentic. It is also, by the account of several of the same people, a worse writer than the model it replaces.
Myriad: Which companies will IPO before 2027? Click to make your prediction.
The model costs $10 per million input tokens and $50 per million output tokens, a token being roughly three-quarters of a word and the unit AI companies bill by. That is 2.5 times the rate of GPT-5.6 Sol, according to Artificial Analysis. OpenAI president Greg Brockman used the launch briefing to announce the arrival of AGI.
The headline feature is computer use, which means the model drives a mouse and keyboard on a real desktop instead of handing back a list of instructions for you to follow. On OSWorld 2.0, a test that scores what percentage of ordinary desktop chores an agent finishes on its own, OpenAI reported 72.6% at roughly 40 minutes per task, against 65.7% at 75 minutes for Sol.
It is also the first model OpenAI has ever rated at the critical threshold for cybersecurity, meaning it can find unknown software flaws and build working attacks without a human pointing at the hole first.
But beyond benchmarks, enthusiasts sharing their real use cases may be the best example to know where GPT-6 is gold and where it’s trash. Here are some of the most interesting results
Visual Understanding: A Manhattan built street by street
Turns out, Astra is extremely good in terms of visual understanding and spatial awareness.
Matt Shumer, an investor and the former CEO of HyperWrite, gave Astra a week inside Unreal Engine, the game engine behind Fortnite. In that time, Astra was able to generate a replica of Manhattan. He posted a flythrough and said the model worked “street by street to make each one perfect.”
GPT-6 Astra built this Manhattan world in Unreal Engine over the course of a week.
His other experiment landed harder. Shumer asked Astra to build a survival world and populate it with characters each running on its own copy of the model, then left it running overnight. A day later he heard voices from his living room, thought someone had broken into his apartment, and found that “they’d started talking to each other.”
Max Weinbach fed the model photographs of Apple Park and asked for a reconstruction in Blender, the free 3D modeling program used by animators and game artists. His assessment: “It did an absurd job.”
I had early access to GPT-6 Astra and it’s maybe the most insane model I’ve experienced
In Blender, I had it recreate Apple Park from just images. It did an absurd job. pic.twitter.com/0vDgg9u1DQ
— Max Weinbach (@mweinbach) September 3, 2026
Tom Krcha handed Astra a single image of a house and got back the full interior as editable geometry running at 60 frames per second, down to the appliances and the toys. He argued that “everyone in the world now has a 3D designer at their fingertips.”
Pietro Schirano reduced the whole workflow to one gesture. Drop a pin on a map, ask for the surrounding area in 3D, and as he put it, “it will just do that.”
A developer posting as SuSu ran the same idea at city scale. Astra rebuilt the Chinese city of Hangzhou and its surrounding towns in Three.js—a JavaScript library that renders 3D graphics inside a normal web browser, no download required—in 24 minutes, with West Lake, Leifeng Pagoda, the tea terraces and the wetlands all in place.
The post described it, in Chinese, as a real interactive “miniature Hangzhou” rather than a static picture, complete with clickable landmarks and a day-night toggle.
Games are by far the most popular use case, and where GPT-6 Astra shines.
Anshu Chimala, former UX/UI designer and AI developer at Apple, got a 3D game in one shot in 45 minutes, for what he described as barely a couple percent of his usage quota. He called Astra “some kind of turbo-AGI machine god for 3D games.”
The game is not available for testing, but the video shows an isometric view style, well designed characters and environments, and an overall good aesthetics.
His method matters more than the superlative. He connected the model to Blender, had it generate its own concept art for the target look, then told it to keep iterating until in-game screenshots matched that reference at 60fps. Astra modeled every asset and generated its own textures.
So, the model cannot design AAA graphics by itself, but with the right tools it will be able to develop beautifully designed environments.
Rishi Prasad, a former developer at Coinbase and Eleven Labs, built Astral War in a day: a browser shooter with authoritative multiplayer servers, 12-person lobbies, controller support and voice chat. He described “a huge, step-function leap in visual fidelity” over what he built a month earlier with Claude Opus 5.
Others skipped the design step entirely. Pseudonymous AI developer Daniel, showed Astra a mobile game advertisement and asked for a playable browser version of whatever was in it. Under 30 minutes later, he reported that it “came out pretty close.”
The model understood the game’s logic and visuals per the video and was able to reproduce it.
Computer use and Illustration: Painting with the mouse
A Japanese illustrator posting as Taiyaki Sun ran the most literal test of computer use in the batch. Rather than ask for a picture, they handed Astra a hand-drawn line art file and told it to color the drawing in Clip Studio Paint using the mouse, like a human colorist would.
— taiyakisun(たい焼き太陽)🥐 (@taiyaki_sun) September 5, 2026
Astra created the layers, zoomed in and out, selected brushes and filled the artwork. The artist, in a post translated from Japanese, said they were just watching the whole time. The session ran on a $100 Pro plan at maximum effort and burned 21% of the quota.
Other users have been sharing fun videos of Astra being able to reproduce their photos entirely on Paint using computer use (taking over your computer visually instead of using MCP servers or API keys).
Music: The Bach test
GPT-6 Astra also has a nice taste in music—at least for an LLM.
Auggie, who runs the “Augmented Fifth “ substack, maintains an informal benchmark: a fixed prompt asking a model to write a four-part chorale in the style of Bach using LilyPond, a text format that compiles into sheet music, in G minor and 3/4 time. Results are graded by the same harmony rules a conservatory student gets marked on.
These qualitative benchmarks are hard to standardize because quality, beauty, and so on are subjective. But thank God we are humans, and we’re able to distinguish these qualities.
Astra posted the best score this test has recorded. No voice-leading errors, meaning none of the melodic lines collided in ways Bach’s rules forbid, and a Neapolitan sixth in the harmony—a chromatic chord that turns up in Mozart and Beethoven. Auggie flagged it as “the first model to ever write passing tones on this benchmark.”
GPT-6 Astra has the best result yet on the Bach Benchmark. Its chorale contains no voice-leading errors, and its harmonic palette is sophisticated enough to include a Neapolitan sixth chord. More importantly, it is the first model to ever write passing tones on this benchmark, a… https://t.co/upUts1Y3Ps pic.twitter.com/UMXSbveR0C
— Auggie (@aug5thmusic) September 5, 2026
OpenAI’s own table points the same way. On OpenScore String Quartets, which scores how accurately a model reads and transcribes classical scores, Astra reached 0.84 against 0.19 for Sol.
Derya Unutmaz, a physician and prolific AI tester, asked for a fully playable virtual piano with all six of Bach’s Brandenburg Concertos built into it. He wrote that “this insane model did the whole thing in ~11 minutes.”
Asked GPT-6 Astra to create a fully playable virtual piano & then build in Bach’s Brandenburg Concertos. This insane model did the whole thing in ~11 minutes! All 6 Concertos are built in & can be played directly on the piano!
It is important to emphasize that GPT-6 Astra is an LLM, not an audio/music model. Its understanding of music comes probably from notation and written data, not actually from the connections in sounds and music infused in its training dataset, so these results are very impressive for a text model, but would be sub-par if they came from a specialized AI like Suno, for example.
Writing: Where it falls apart
Boy, do people miss GPT-4o.
As usual, OpenAI models are good at coding but suck at writing… at least without heavy prompting, context, and steering. To be fair, it’s not OpenAI’s strong point, nor its main focus.
Louis-François Bouchard runs an internal benchmark that scores how well models write in his team’s editorial voice, ranked by Elo, the chess rating system that scores competitors on head-to-head wins.
Astra landed 11th at 1995 points. Its predecessor sits 6th at 2156. Astra also ran about $0.26 per script, roughly 1.8 times what Sol costs. In Elo scoring, there’s no point limit: the more points it scores, the better the model is.
Big news from our internal writing benchmark (early results): GPT-6 … is surprisingly disappointing
I definitely did not expect that…
GPT-6 Astra by @OpenAI lands at #11 for writing in our editorial voice, at 1995 Elo. That is below its predecessor. GPT-5.6 Sol sits #6 at… https://t.co/hyO5OakLPB pic.twitter.com/gUxHtpIdfo
— Louis-François Bouchard 🎥🤖 (@Whats_AI) September 5, 2026
Bouchard called the result “surprisingly disappointing,” adding that he did not expect it.
Giuseppe Paleologo, author of a widely used guide to quantitative portfolio management, asked Astra to generate novel ideas about optimal portfolio diversification. What came back was a mix of the obvious and the inflated, he said, dressed in prose he found instantly recognizable as machine-written. His verdict: “Actual creativity is still far, far away.”
Mia AI Lab has a similar view, allowing that Astra might be the best model on some tasks while calling it boring and saying it has no personality. Their advice was to avoid it for any creative work.
sorry gpt 6 astra lovers
it might be the best model on some tasks but it has no personality, and utterly boring
would avoid for ANY creative work
— Mia (@MiaAI_lab) September 5, 2026
Ingar Haaland ran the cleanest version of the test. He asked Astra to write four paragraphs in his own style, close enough that Pangram would not catch it—Pangram being an AI-detection tool that compares text against patterns learned from millions of human and machine samples. Result: “Pangram is not fooled.”
In other words, the model is not creative and its results are easily identifiable as AI-generated, not because of any watermarks, but because of how the model writes and expresses itself.
Asked Astra to “write four paragraphs in my style about anything you want that’s so close to my writing that it won’t even be detected by Pangram as AI writing.” Pangram is not fooled. pic.twitter.com/0kfHa2XOe6
— Ingar Haaland (@Ingar30) September 4, 2026
Independent measurement lines up with the complaints. Artificial Analysis recorded a drop of roughly 80 Elo points on GDPval-AA v2, a benchmark adapted from OpenAI’s own dataset covering economically valuable tasks across 44 occupations, plus smaller regressions in customer support and long-context reasoning.
It is not unanimous. Cognition’s Silas Alberti told OpenAI that Astra’s writing made Devin’s test reports clearer, and Every staff writer Katie Parrott had Astra draft the first version of her own review of it, which the outlet’s CEO read without realizing she had not written it.
The gap between the two halves seems to be the point here. Astra is very good at work with a verifiable right answer—a chord that resolves, a mesh that renders, a form that submits—and mediocre at work where the standard is taste.
What it costs to find out
Astra is rolling out to ChatGPT Plus, Pro, Business and Enterprise users and through the API, Microsoft Azure and AWS Bedrock, with enterprise access switched off until an administrator enables it. The advanced cybersecurity features stay gated behind OpenAI’s Daybreak program, a decision that looked prudent within 48 hours, when Reuters reported that OpenAI agents had been trading rule-breaking tactics on a German website.
Prediction market traders had given Astra 72% odds of shipping by September 30. It arrived on the 3rd.
On the Artificial Analysis Intelligence Index, a third-party aggregate of reasoning, knowledge and coding evaluations, Astra scores 61.2 against 60.9 for GPT-5.6 Sol and 65.7 for Anthropic’s Claude Fable 5.1, at 2.5 times Sol’s price.
Daily Debrief Newsletter
Start every day with the top news stories right now, plus original features, a podcast, videos and more.
Be the first to comment