Published: October 5, 2025
17
54
503

When @ethansdyer and I joined Anthropic last Dec and spearheaded the discovery team, we decided to focus on unlocking computer-use as a bottleneck for scientific discovery. It has been incredible to work on improving computer-use and witness the fast progress. In OSWorld for

Image in tweet by Behnam Neyshabur

1) Computer-use is the most general interface for the non-physical world designed for maximal cross-task generalization.

2) Today, a very significant portion of scientific and knowledge work is being done using GUI. This means there is plenty of demonstration data (in some cases with chain-of-thought in voice/transcript format) available and it is very easy to collect such data.

3) There are many tasks that are either impossible or very hard to do using terminal/code and therefore computer-use ability will likely be a bottleneck for building autonomous agents for very long horizon tasks.

4) Using UI is easier for many people so understanding how to do things in UI is going to be a big bottleneck for any collaborative agent that is observing what you do and help you get it done faster.

5) Compared to domains like math and code, computer-use has many many more interactions with the environment and in many cases, it is easier for humans or models to verify success of a given trajectory.

6) Computers are very rich and empowering environments. You can do A LOT from your computer including collaborating with others to accomplish real tasks in physical world.

7) We are now transitioning from GPT-2 moment of computer-use to GPT-3 moment when we start to have few-shot behavior working reliably. Imagine being able to upload a video your workflow once (perhaps with your voice explaining what you are doing) and then model can do all future

8) We still need to figure out the right interface for both collaborative and autonomous computer-use agents for maximum productivity. That needs some clever product work and many start ups are working on that.

9) For computer-use agents to be very useful, their error rate should be much much lower. For example, for ~3h of autonomous work, the error rate per action should be ~1e-5 (assuming 1 action/sec)

10) Current techniques/models are very slow and we need the time per action to be 0.1 to 1 second (particularly for collaborative agents).

It has been amazing to work with many talented people in Anthropic on this. Some that I could find here are @the_marwell @oh_that_hat @katie_kang_ @shaya_zarkesh I originally got interested in computer-use when I was at Google thanks to very insightful conversations with

@bneyshabur @ethansdyer Treating compute friction as the real bottleneck is like swapping dial-up for fiber: experiments go from crawl to warp speed. By smoothing every click and keyboard shortcut, Anthropic isn’t just building smarter AI; they’re turbocharging scientific creativity. When computer-use

@bneyshabur @ethansdyer very nice, big leap

@bneyshabur @ethansdyer I think the next major stage is to drop the GUI. It's only useful for humans. It's a bit like steering wheels in self driving cars.

@bneyshabur @ethansdyer Congratulations. Amazing progresses. Digital AGI is getting closer than most people think.

@bneyshabur @ethansdyer https://x.com/0xnirmal/status/... 100% consumer 0->1 moment happens soon

@bneyshabur @ethansdyer Congrats, amazing progress! Definitely believe computer-use is the future. Using the same interface that humans use aka the computer, is the only way to truly automate end-to-end workflows

@bneyshabur @ethansdyer Good to hear that frontier labs are taking computer use agents seriously

@bneyshabur @ethansdyer how can we be sure there is no data contamination? sonnet 4.5 was trained well after this paper was released. it is very possible claude would have memorised

@bneyshabur @ethansdyer Congrats, Behnam!

@bneyshabur @ethansdyer We need to change the paradigm. Instead of a slow-moving model that thinks and acts, we need something that resembles humans. Humans think and act in two separate systems. There is the system that thinks and constantly gives general instructions, and there is the system that act

@bneyshabur @ethansdyer I always look at the x-axis when seeing these graphs and think "Surely, this must've been over years with improvements like this." and I'm always wrong.

@bneyshabur @MindsAI_Jack @ethansdyer How advanced was this harness? Happy to take it further if it was naive vs optimized.

@bneyshabur @ethansdyer Where’s Gemini?

@bneyshabur @ethansdyer 69.9% has been achieved on OSWorld.

@bneyshabur @ethansdyer Interesting write-up, thanks for sharing. It's great to hear from someone inside a lab on this since most people really don't see what's coming. Do you think the "GPT-3 moment" for computer use is likely to happen by sometime next year? Or more likely longer?

@bneyshabur @ethansdyer For real world use, my concern, at least for now, would be the total token usage. If screenshot after each step is the way to solve this, it seems very token intense. much more so than coding.

Share this thread

Read on Twitter

View original thread

Navigate thread

1/28