GPT-6 Astra Just Launched With “Computer Use” – Here’s What Actually Matters for the Rest of Us

GPT-6 Astra brings computer-use AI to OpenAI’s flagship model, but limited access, unknown pricing, and missing independent testing demand caution.

OpenAI released GPT-6 Astra today. The headline feature isn’t a smarter chatbot. It’s a model that can operate a computer the way a person does — clicking through menus, filling out forms, moving across web pages. I’ve tested AI tools since the GPT-3 beta days. This is the first release in a while that made me stop and think hard about my workflow a year from now.

Here’s what’s real, what’s still vapor, and why it matters if you run marketing operations instead of just watching AI headlines.

Why It Matters

“Computer use” isn’t a new idea. Anthropic ran a beta version back in 2024. Perplexity shipped something similar earlier this year. What’s different is that OpenAI now builds it into its flagship model, and the company frames it as a threshold moment. OpenAI president Greg Brockman told reporters the model can move through spreadsheets and web pages “often at superhuman speed.” He called computer use core to this release, not a side feature.

For marketers, that’s the actual story. If a model can reliably navigate ad platforms, CRMs, and reporting dashboards without a human clicking every button, that changes what “automation” means for a MarTech stack. An AI that drafts your ad copy is one thing. An AI that logs into the platform and launches the campaign is something else entirely.

I want to flag that word “if.” We don’t have independent testing on real-world business software yet. Right now we only have OpenAI’s own demo and benchmark numbers.

Technical Details

Astra is, by OpenAI’s own account, its largest training effort to date. Aidan Clark, the company’s VP of research, said the team pretrained this model on more than 100,000 GPUs at OpenAI’s Stargate site in Texas — a first for the company. That scale matters less to the average marketer than what it enables. Still, it’s worth knowing this wasn’t an incremental update.

Astra also crosses OpenAI’s internal “critical cybersecurity capability threshold.” Under certain conditions, the model can find and exploit unknown security flaws without a human in the loop. That’s a meaningful capability jump, and it explains why OpenAI is staging the rollout instead of releasing it to everyone at once.

Performance & Evidence

OpenAI’s own benchmarks put Astra well ahead of its predecessor, GPT-5.6 Sol, and ahead of Anthropic’s current top models. On ARC-AGI-3 — a benchmark that tests reasoning on novel problems rather than memorized patterns — OpenAI reported a score of 98.6% for Astra. GPT-5.6 Sol scored only 7.8%, and Claude Opus 5 scored 30%. On a cybersecurity benchmark called ExploitBench, OpenAI says Astra hit 100%, compared to 78.5% for GPT-5.6 Sol.

Those numbers look dramatic. But I’ll say what I say about every vendor-reported benchmark: early results suggest something significant, yet a company grading its own model against its own prior model isn’t the same as independent verification. ARC-AGI-3 is a genuinely hard benchmark, so a jump this large deserves attention. I still wouldn’t build a client strategy around it until third-party testing confirms it holds up on messier, real-world tasks.

Pricing & Availability

Here’s where I’d tell any client to pump the brakes. Astra isn’t rolling out broadly yet. Right now access is limited to select enterprise customers, including participants in OpenAI’s cybersecurity-focused Daybreak program. OpenAI says general availability for Plus, Pro, and Enterprise users — plus API and AWS access — is coming “in the coming days.” The company hasn’t set a firm date, and it hasn’t published pricing.

That “contact us for access” pattern is one of my long-standing pet peeves in this industry, and it’s happening again here. Nobody can plan a budget around this model until OpenAI publishes pricing.

One more limitation worth flagging: the public version will refuse advanced cybersecurity tasks even after general release. The full-capability version stays restricted to vetted partners.

Industry Implications

Computer-use agents could matter more to marketing teams than another chatbot upgrade, assuming they mature the way OpenAI is betting they will. Think about the grunt work inside a MarTech stack: pulling reports from five different dashboards, updating spend allocations across ad platforms, reconciling data between a CRM and an email tool. This kind of multi-step, click-heavy work has resisted automation for years because it requires navigating inconsistent, non-API-friendly interfaces.

I’ve built AI-assisted workflows that boosted client output significantly, but almost all of that gain came from tools with clean APIs. A model that can operate any interface — including ones without APIs — would remove a bottleneck I hit constantly. I’m not ready to call it solved. I am saying this is the first version of this idea a major lab has positioned as a flagship feature rather than a research demo.

Limitations & Open Questions

A few things I’d want answered before I recommend this to a client:

  • How does it perform on real, messy enterprise software rather than staged demos — with login walls, CAPTCHAs, and the inconsistent UI that actual marketing tools have?
  • What happens when it clicks the wrong thing in a live campaign environment? Nobody wants an agent that mis-clicks its way into pausing a client’s ad account.
  • What will it actually cost per task or per seat once OpenAI announces pricing?
  • How will OpenAI’s staged safety review hold up under outside scrutiny? The company ran this review through a voluntary government framework, and it hasn’t published the details. Brockman also declined to detail the process when reporters pressed him.

We still need independent benchmarks and real production testing here. Vendor demos are choreographed by definition.

GPT-6 Astra benchmark dashboard showing computer-use automation, AI performance, enterprise access, and safety limitations

What This Means for Developers and Marketers

Developers building on GPT-6 Astra face a practical integration question. Computer-use agents that navigate arbitrary interfaces can reduce the need for custom API connectors, but they also introduce new failure modes that traditional API calls don’t have. For marketers and MarTech buyers, the honest answer right now is simple: watch, don’t buy. Access stays limited, pricing isn’t public, and independent real-world testing hasn’t closed the gap with OpenAI’s benchmark numbers.

The Bigger Picture

You don’t have to buy Brockman’s framing that this marks the start of an “AGI era” — a term AI researchers still can’t fully agree on — to see the broader trend. Every major lab now races toward agents that act inside software instead of just generating text about it. Anthropic got there first with a beta. Perplexity built its own version. Now OpenAI ships one as a flagship capability.

The real question for anyone running a marketing or ops team isn’t whether this looks impressive in a demo. It’s whether it holds up on a Tuesday afternoon, inside your actual ad platform, with your actual account structure, without a human babysitting every click. That’s the test I’ll watch for once broader access opens up. I’d encourage anyone evaluating this for their stack to wait for that evidence before rearranging a budget around it.