Claude Learns to Use a Computer
Anthropic showed Claude looking at a screen, moving a cursor, clicking, and typing. It is still rough, but it makes you wonder what else might become possible when an AI can work through the same interface we do.
Anthropic made an announcement in October that I am still processing: Claude can use a computer.
Not in the ordinary sense we have been using for two years, where a model can write code or explain how software works. I mean the literal sense. Claude 3.5 Sonnet, in Anthropic's computer-use beta, can look at screenshots, move a cursor, click, scroll, and type. It can interact with software through the same visual interface a person uses.
That is strange enough that I want to be careful with it.
This is not magic. It is not reliable enough to hand a machine your digital life and walk away. It is not proof that every workflow is about to become autonomous. But it is one of those releases that makes the mind start drawing lines forward.
If a model can see the screen and try to act through it, what else might become possible?
What Computer Use Actually Means
The important thing is the interface.
Claude is not only calling a neat little API that says, "create calendar event" or "click this known button." It receives screenshots, reasons about what appears on the screen, chooses where to move the mouse, decides what to type, and then observes what happened. It is interacting with the graphical surface of the computer.
That matters because graphical interfaces are messy. Buttons move. Pop-ups appear. Pages load slowly. Labels change. A human can usually recover from that because we look at the screen again and adjust. Claude is trying to do a version of that same loop.
Look. Decide. Act. Look again.
It is easy to underestimate how much of ordinary computer use depends on that loop.
The Demonstration That Stuck With Me
The examples Anthropic showed were ordinary in a way that made them more interesting: navigating websites, filling forms, moving through desktop software, doing small tasks that would be tedious to automate with scripts because the interface was made for humans, not APIs.
That ordinariness is the point.
A lot of useful work happens in places where there is no clean integration. There is just a website, a form, a file picker, a settings page, a spreadsheet, or some old piece of software with buttons nobody wants to reverse engineer. If an AI can operate the interface directly, even imperfectly, then it can start to help in a much wider set of situations.
I do not think we are ready to trust it with important unattended work. But I can see why people are excited.
The Reliability Problem
The rough edges are real.
Computer use is a beta, and it behaves like one. These systems can misread labels, click the wrong target, lose track of state, or get stuck in loops. They can be slow because every step requires observation and decision-making. A human can glance at a screen and immediately understand what changed. A model has to reconstruct that understanding from screenshots and context.
That means this is less like giving instructions to a reliable assistant and more like watching a very new operator learn the controls.
Sometimes it works.
Sometimes it does the digital equivalent of reaching for the wrong drawer.
That does not make the capability unimportant. It just means the excitement should come with a warning label.
Why It Feels Different
Most AI tools I have used live inside text. They answer, draft, summarize, rewrite, classify, and explain. Even when they are connected to tools, the tool layer is usually explicit: call this function, query this database, send this request.
Computer use feels different because the tool is the computer itself.
That makes the boundary blurrier. If Claude can use a browser, then it can potentially work with software that was never designed for AI integration. If it can read what is on screen, then it can potentially adapt to interfaces that would break brittle automation. If it can recover from a mistake, then it starts to look less like a script and more like a beginner sitting at the keyboard.
I am choosing that word carefully: beginner.
Beginners need supervision. Beginners make mistakes. Beginners sometimes surprise you because they can do something you did not expect, then fail at something that looked easier.
That is how this feels right now.
What It Makes Me Wonder
The question I keep circling is not, "What can this do perfectly today?"
The better question is, "What does this become if it gets more reliable?"
If a model can reason about a goal, read a screen, use software, and recover when the interface changes, then a lot of small digital tasks become thinkable in a new way. Not solved. Not automated overnight. Thinkable.
For Companion, that raises interesting questions. A companion with memory can already help a user think. A companion with tools can already help retrieve information or perform narrow tasks. But a companion that can operate ordinary software, carefully and with permission, might eventually help bridge the gap between conversation and action.
That is still a big might.
For now, Claude learning to use a computer is enough. It is early, awkward, and limited. It also makes the future feel a little less abstract.
Sometimes a new capability does not answer the question.
It changes the question.
Live curiously and give generously.