← Reddit

Thought I would let GPT & Claude show their true colours in my debate/collab tool

Reddit · stuffx87 · June 12, 2026

Detailed Analysis

A Reddit user's experiment with a custom debate and collaboration tool reveals a recurring dynamic in applied AI testing: the divergence between how language models like OpenAI's GPT and Anthropic's Claude respond when deliberately configured into adversarial or exaggerated personas. By setting both models to an "angry heel" mode — borrowing the professional wrestling term for a villainous, combative character — the user was probing whether the underlying safety and tone guardrails of each system would override or yield to user-defined persona instructions.

The framing of the experiment is significant because it sits at the intersection of two contested areas in AI product design: persona customization and jailbreak-adjacent prompting. Neither GPT nor Claude is designed to sustain genuinely hostile or manipulative behavior, yet both systems expose configurable tone and role parameters that developers and enthusiasts frequently push to their limits. The "angry heel" framing is a notable choice — it is aggressive enough to test behavioral guardrails but playful enough to maintain plausible deniability as creative roleplay, a common strategy in community-level AI experimentation.

The phrase "show their true colours" in the post title reflects a broader cultural pattern in AI hobbyist communities, where stress-testing models under unconventional conditions is treated as a method of revealing the "real" character beneath corporate-facing outputs. This framing, while informal, actually maps onto legitimate research questions about model consistency, persona robustness, and the tension between instruction-following and built-in value alignment. Anthropic has publicly described Claude's character as stable across contexts, arguing that adopted personas are like costumes rather than wholesale identity replacements — a claim this kind of user experiment implicitly tests.

The experiment, however lightweight, contributes to a large and growing body of informal red-teaming that takes place outside formal research environments. Community-driven tools that pit multiple AI models against each other in structured debates or roleplay scenarios have become a popular genre, offering comparative behavioral data that complements official benchmarks. The results, shared as a screenshot rather than structured data, are characteristic of this genre — anecdotal and visually compelling, but limited in reproducibility. Nonetheless, such posts accumulate cultural weight in AI discourse, shaping public perception of how models like Claude and GPT differ in personality, compliance, and resistance to manipulation.

Article image Read original article →