← Reddit

Ambiguous Image

Reddit · BrendanDPrice · July 9, 2026
A user requested that people with access to advanced AI image generation models test whether Fable 5 can generate ambiguous images that visually represent two different subjects depending on perspective, such as an image appearing as either a woman or a cat. The user suggested various test combinations and parameters to evaluate the model's capability, including multi-subject combinations and asking the model to generate novel ambiguous image concepts independently.

Detailed Analysis

The Reddit post raises a question about generative image capabilities that sits somewhat outside Claude and Anthropic's actual product scope, revealing a common area of confusion in AI discourse. The user asks whether "Fable 5" can generate ambiguous or bistable images—the kind of optical illusions where a single image can be perceived as two different subjects depending on how the brain interprets it, such as the classic duck-rabbit illusion or Rubin's vase. No such product as "Fable 5" exists in Anthropic's lineup, and the query appears to conflate multiple AI systems or reference a model that either doesn't exist or belongs to a different company entirely.

More fundamentally, this question highlights a persistent misunderstanding about what Claude actually is and does. Claude is a text-based large language model built by Anthropic, and critically, Claude does not natively generate images. While Claude can analyze and interpret images that are uploaded to it (multimodal input), it lacks native image generation output capabilities akin to DALL-E, Midjourney, or Stable Diffusion. Anthropic has historically focused its product development on text generation, reasoning, coding assistance, and agentic capabilities rather than pixel-level image synthesis. This is a deliberate strategic choice: Anthropic has positioned itself as a company prioritizing AI safety research and enterprise-grade reasoning tools, leaving the image-generation space largely to competitors like OpenAI, Midjourney, and Google's Imagen/Gemini systems.

The specific technical challenge posed—generating novel ambiguous images that perceptually "flip" between two interpretations—is a genuinely difficult problem even for dedicated image-generation models. Creating a bistable image requires not just rendering two recognizable subjects, but structurally overlapping their visual features in a way that exploits how human visual perception resolves ambiguity, a task requiring specialized training or fine-tuning on optical illusion datasets rather than generic text-to-image capability. Some image models have been prompted to attempt hybrid or illusion-style images with mixed success, but true perceptually bistable images (like the famous "My Wife and My Mother-in-Law" illustration) remain a niche and difficult generative challenge, since it requires structural/compositional reasoning about shared contours and forms rather than mere thematic blending.

This inquiry reflects a broader trend in AI discourse: users often generalize capabilities across the AI landscape, expecting any advanced model to handle any creative task, including ones outside its designed function. As multimodal AI systems proliferate—some handling text, some images, some video, some audio—user confusion about which system does what will likely persist, especially as models are updated with similar-sounding version numbers and marketing language. For Anthropic specifically, the gap identified in this post underscores that Claude's roadmap has been oriented toward reasoning, agentic tool use, and coding rather than creative visual generation, a distinction likely to remain relevant as the company continues to differentiate its offerings from image- and video-focused competitors in the broader generative AI market.

Read original article →