← Reddit

I played an entire game of Sorry with Claude.

Reddit · colt1776 · August 16, 2026

Detailed Analysis

A Reddit user's lighthearted experiment—playing a full game of the classic board game Sorry! against Claude—offers a small but telling glimpse into how everyday users are probing the boundaries of large language models' capabilities outside of coding, writing, or research tasks. Rather than relying on a digital interface or plugin, the user took photographs of the physical board after each turn and fed them to Claude, asking it to interpret the game state, track pieces, and make strategic decisions. The informal, almost playful nature of the post—posted to r/ClaudeAI with a "guess who won?" framing—underscores how much grassroots experimentation is happening among Claude's user base, often disconnected from Anthropic's official marketing or enterprise use cases.

This kind of test matters because it stresses multimodal reasoning in a domain LLMs weren't explicitly trained for: interpreting photographs of a physical game board, maintaining state across multiple turns, and applying rules-based strategic reasoning in a visually grounded, real-world context. Board games like Sorry! require tracking piece positions, understanding turn-based mechanics, interpreting card draws, and making tactical decisions (like when to bump an opponent's piece)—all from images that may vary in angle, lighting, and clarity. Success or failure here reveals a lot about how well vision-language models generalize from static image understanding to dynamic, sequential state-tracking, a capability with implications far beyond games, including robotics, physical-world assistants, and accessibility tools for people who might photograph documents, diagrams, or environments for an AI to reason about over time.

The exercise also reflects a broader trend of users treating frontier AI models as general-purpose reasoning partners rather than narrow tools. Where earlier chatbot interactions were largely confined to text-based Q&A, users are increasingly testing multimodal models like Claude (which supports image inputs) in embodied, real-world scenarios—cooking from a photo of ingredients, diagnosing a plant's health from a leaf picture, or, in this case, playing a physical board game turn by turn. These informal stress tests, while anecdotal, often surface capability gaps or surprising strengths well before formal benchmarks do, and they shape public perception of what "AI that reasons" actually looks like in practice.

Finally, this kind of casual, community-driven experimentation is part of a larger cultural pattern around Claude specifically. Anthropic has cultivated a reputation for Claude's conversational quality, personality, and willingness to engage in creative or playful exchanges, and communities like r/ClaudeAI have become informal testing grounds where users share screenshots, transcripts, and quirky use cases. Posts like this one, even without deep technical rigor, contribute to a body of anecdotal evidence that shapes user trust, expectations, and enthusiasm—serving as a form of organic, decentralized evaluation that complements (and sometimes outpaces) formal benchmarking efforts from AI labs themselves.

Read original article →