AI Model Reviews

I Pushed GPT-6 Astra Computer Use to Its Limits

Direct Answer

GPT-6 Astra computer use is already useful for delegated visual work, but it is not a universal replacement for skilled operators or direct software integrations. Futurepedia's ten-app test found three good fits: completing a task inside unfamiliar software, handling repetitive setup inside a familiar tool, and producing an editable first pass while the human works elsewhere.

The failures are just as instructive. Forced mouse-and-keyboard control performed poorly on a complex Illustrator recreation and took almost twice as long as Python in Blender. Computer use is the fallback for software that exposes no better interface, not the default merely because watching an agent move a cursor looks impressive.

The practical rule: delegate outcomes, not clicks. Give the agent the goal, quality bar, allowed tools, maximum iterations, and approval boundaries. Let it use a direct API, script, or application connector when that path is more reliable.

Watch the Ten-App Test

Credit and evidence note: application results, timings, prompts, and judgments come from Futurepedia's creator test, published 15 September 2026. The video deliberately blocks scripts, APIs, and MCP tools in several trials to isolate cursor-based computer use. That makes the comparison useful, but unlike a normal production workflow.

The Ten-App Scorecard

ApplicationTaskReported resultEditorial read
Google Earth StudioAnimate a Utah fly-throughUsable after three follow-upsStrong fit for unfamiliar visual software
Canva DrawCreate original drawingsIncreasing detail over three attemptsCapable, but creative variety stayed narrow
PhotoshopBuild a layered thumbnailEditable result after about one hourUseful background delegation, slower than the expert
Premiere ProAdd text, motion, and sound effectsGood final sequence after about 25 minutesStrong for bounded, tedious edits
IllustratorRecreate a complex vector imagePoor through UI; better with PythonWrong task for cursor-only control
BrickLink StudioModel Delicate Arch in LEGO257 pieces and 27 build stepsConvincing first-pass structured output
OpenRocketDesign and simulate a stable rocketIterated until constraints passedPromising, but expert validation still required
AlgodooBuild a physics chain reactionCompleted, visually modestSuccessful but not a demanding proof
BlenderRecreate Delicate ArchUI: ~105 min; Python: ~54 min and betterDirect control surface clearly wins
GarageBandPlay and arrange a seven-track songFinished in 52 minutesResourceful around input limits, musically basic

These are demonstrations by one creator, not standardized benchmarks. The creator accelerated the footage, used different prompt depths, and sometimes had prior familiarity with the software. The useful evidence is the pattern across tasks, not a universal ranking.

Creative Software: Good Assistant, Uneven Art Director

Google Earth Studio produced the clearest end-to-end win. The first pass took 11 minutes 49 seconds and established the route, camera, and keyframes. Three follow-ups added roughly 23 minutes and corrected pacing, direction, and the approach to Delicate Arch. For someone who uses the software infrequently, a usable result in about 35 minutes can be valuable even when the agent is not faster than an expert.

Canva revealed a different limitation. The agent spent under four minutes on the first drawing, 7 minutes 14 seconds on the second, and 20 minutes 47 seconds on the most detailed version. Complexity increased, but each concept repeated a similar animal-carrying-a-structure motif. Longer work did not automatically produce broader creative exploration.

Photoshop took roughly an hour across three rounds. The final file retained editable layers, which matters more than a flattened preview, but the creator could have assembled it faster himself. The gain was asynchronous labor: he reviewed results from his phone while the agent worked on a separate monitor.

Premiere was the most practical Adobe result. Astra added text behind the subject, animated placed images, searched a sound-effects project, and synchronized the choices. Twelve minutes for the first pass plus 13 minutes for the revision is reasonable for a bounded edit that would otherwise require manual keyframing.

Illustrator exposed the ceiling. The forced interface run spent 40 minutes on a poor first pass and another 52 minutes without approaching the reference. When allowed to use Python, Astra made a better version in 20 minutes and improved it over two further passes. Neither matched the target, but the direct method was plainly more effective.

LEGO, Rockets, and Physics: Completion Needs Validation

BrickLink Studio was a striking first try: Astra represented Delicate Arch with 257 pieces and organized the model into 27 assembly steps. That demonstrates competent manipulation of an unfamiliar structured design tool. It does not prove that every connection is physically robust or that the final parts list is economical.

In OpenRocket, the agent iterated until its simulated design exceeded 1,000 meters and met the recovery constraint. That is a successful software workflow, not an engineering certification. Simulation inputs, component availability, stability margins, launch conditions, and applicable safety rules still need review by someone qualified.

The Algodoo chain reaction completed the requested sequence but did not look especially difficult. This is a useful reminder to set a quality bar before starting. “The simulation runs” and “the design is elegant, robust, and worth keeping” are different acceptance tests.

Blender Proves Why the Control Surface Matters

The cursor-only Blender attempt began with a weak result after 15 minutes. Astra researched reference images, rebuilt the form, adjusted materials through the node editor, inspected screenshots, and improved the surrounding scene. After roughly one hour 45 minutes, it reached a respectable but still intermediate-looking result.

The unrestricted run used Blender's Python interface and produced a visibly better scene in about 54 minutes. The comparison does not show that computer use failed. It shows that visual understanding and direct structured control work best together. Cursor control is valuable for surfaces that have no reliable API; it should not replace a better interface to make the automation feel more human.

Loop with a budget: self-critique can improve an artifact, but a weak capability can also consume hours chasing an unreachable quality bar. Set a maximum number of iterations, elapsed-time ceiling, or usage budget, then escalate to human review.

GarageBand Shows Adaptation, Not Musicianship

Using iPhone Mirroring, Astra tapped virtual drums and instruments rather than arranging notes directly in a desktop timeline. It could not hear the metronome in real time, press multiple points simultaneously, or reliably hold a note. It compensated by using quantization, splitting drums across tracks, and toggling sustain to extend notes.

That adaptation is the interesting capability. The final seven-track song took 52 minutes and was musically basic, but the agent identified input constraints and changed strategy without being explicitly told how. A future workflow should specify genre, tempo, reference tracks, instrumentation, song structure, and an evaluation checklist rather than simply asking for something “more complex.”

Choose the Best Authorized Control Surface

Control surfaceUse it whenMain tradeoff
Computer useThe task is visual, the software is unfamiliar, or no integration existsSlower and sensitive to layout or UI changes
API or connectorThe action is structured and repeatableRequires scoped credentials and may not expose every feature
Script or application APIPrecision, iteration, and artifact structure matterNeeds technical review and sandboxing
Human operatorTaste, expert judgment, liability, or irreversible consequences dominateHuman attention is scarce and synchronous

OpenAI describes Astra as its strongest computer-use model and reports better performance and lower simulated task time than GPT-5.6 Sol on OSWorld 2.0. Those measurements support the direction of improvement; they do not guarantee success in a specific application. Evaluate on the software, files, and acceptance criteria that matter to you.

A Safe First Delegation

  1. Duplicate the source file. Work in a copy of the Premiere project, Photoshop document, Blender scene, or other artifact.
  2. Limit access. Expose only the folders, applications, sites, and accounts needed for the task.
  3. Define “done.” Specify the deliverable, editable format, visual references, constraints, and quality checks.
  4. Choose the method deliberately. Permit the safest effective API or script when appropriate instead of requiring cursor control.
  5. Cap the loop. Set a maximum number of attempts, time limit, or usage budget.
  6. Gate consequences. Require approval before publishing, sending, purchasing, overwriting originals, deleting, or changing permissions.
  7. Inspect the artifact. Review layers, timelines, geometry, formulas, simulation inputs, and exported output, not only a screenshot.

OpenAI's desktop guidance distinguishes the built-in browser from a connected Chrome profile: the former uses its own browser state, while the Chrome route is for existing signed-in sessions and extensions. That difference should influence which accounts and data you expose.

Video Chapters

TimeTopicTimeTopic
00:00Why computer use is the leap10:51BrickLink, OpenRocket, and Algodoo
00:36Google Earth Studio12:43Blender: cursor versus Python
03:01Canva Draw15:51GarageBand through iPhone Mirroring
04:49Adobe Photoshop19:48Final thoughts
06:52Adobe Premiere Pro
08:31Adobe Illustrator

Verdict

Astra's computer use is most valuable as asynchronous, cross-application labor. It can get a non-expert surprisingly far inside an unfamiliar tool and can remove setup or repetitive work from an expert's workflow. It remains slower than a skilled operator on some tasks, and the visual interface can be a severe bottleneck when the software provides a structured control path.

The right adoption question is not “Can it use this app?” It is “Can it return an editable, reviewable result at a lower cost in human attention, inside boundaries we trust?” This test provides several credible yeses, several clear noes, and a much better way to decide.

Sources and Further Reading

Publication date follows the primary video's official YouTube date: 15 September 2026. Editorial review: 17 September 2026. Product access, pricing, usage limits, and interface behavior may change.

Common questions

What was GPT-6 Astra best at in this test?
It was most useful in unfamiliar or tedious visual workflows with a clear finish line. Google Earth Studio reached a usable result after feedback, Premiere handled keyframed effects and sound selection, and BrickLink produced a structured model on the first pass.
Did GPT-6 Astra outperform experienced creative professionals?
No. The creator explicitly concludes that Astra is not better than someone who has spent years mastering Blender, Photoshop, Premiere, or the other tested tools. Its value is reaching competent results in unfamiliar software and removing repetitive steps from familiar workflows.
Should Astra use the interface when an API or scripting option exists?
Usually let it choose the most reliable authorized method. In both Illustrator and Blender, the direct Python route was faster and produced a better result than forced cursor-only computer use.
How long did the tasks take?
They ranged from under four minutes for a simple Canva drawing to roughly 105 minutes for the cursor-only Blender sequence. Many ran in the background, so elapsed time and human attention time were different.
How should I test computer use safely?
Start with a reversible task in a duplicate file or test account. Limit accessible folders and apps, define a stopping condition and maximum iterations, and require review before saving over originals, publishing, purchasing, sending, or deleting.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call