← Google News

Forget about prompts: Anthropic has made it possible to train Claude using screen recordings - Mezha

Google News · July 22, 2026
Forget about prompts: Anthropic has made it possible to train Claude using screen recordings Mezha [truncated: Google News RSS provides only a snippet, not full article

Detailed Analysis

Anthropic has introduced a new capability that allows Claude to learn from screen recordings rather than relying solely on written prompts to understand how a task should be performed. Instead of a user painstakingly describing a workflow in text—clicking through menus, navigating software interfaces, or performing a multi-step business process—the user can simply record their screen while performing the task once, and Claude can ingest that recording to learn the pattern of actions involved. This shifts the paradigm from prompt engineering, where success depends heavily on how precisely a user articulates instructions, toward a more intuitive demonstration-based model of teaching AI systems, akin to how a human trainee might learn by watching a colleague perform a task.

This development matters because it addresses one of the most persistent friction points in deploying AI agents for real-world work: the gap between what a human intends and what they can effectively communicate through text alone. Many business processes are inherently visual and procedural—involving specific sequences of clicks, form fields, and application switches—and these are often difficult to describe accurately in words, especially for non-technical users. By allowing Claude to observe screen recordings, Anthropic is effectively lowering the barrier to customizing AI behavior for specific organizational workflows, reducing the need for specialized prompt-writing skills and making automation more accessible to a broader range of employees and departments. It also suggests progress in Claude's multimodal reasoning capabilities, as the model must interpret sequences of visual information, correlate them with on-screen text and UI elements, and infer the underlying intent and logic of the demonstrated task.

This move fits into a broader industry trend of AI agents evolving from passive chatbots into active participants capable of operating software on a user's behalf—a category often referred to as "computer use" or agentic AI. Anthropic has already been building toward this with earlier releases enabling Claude to control a computer interface directly, taking screenshots, moving a cursor, and clicking through applications. Screen-recording-based training extends this trajectory by adding a learning-from-demonstration layer, which is a technique long used in robotics and reinforcement learning, now being adapted for knowledge-work automation. It positions Claude to compete more directly with other agentic AI efforts from OpenAI, Google, and various enterprise automation startups that are racing to make AI systems capable of executing complex, multi-step digital tasks with minimal human oversight.

Taken together, this feature reflects Anthropic's strategic emphasis on practical enterprise adoption over pure benchmark performance. Rather than simply improving raw language understanding, the company appears focused on reducing the operational overhead of deploying AI in real business contexts—recognizing that many potential users are bottlenecked not by model capability but by the difficulty of specifying what they want the model to do. If screen-recording-based training proves reliable and scalable, it could meaningfully accelerate enterprise adoption of Claude for repetitive digital workflows, from data entry and customer support processes to software testing and back-office operations, while also raising new questions about data privacy, security, and the governance of AI systems that can observe and replicate sensitive on-screen activity.

Read original article →