← Reddit

UPDATE: You asked how the orange negotiation would go against a smaller model. Fable 5 vs Haiku 4.5. It was a massacre.

Reddit · kizmania · June 11, 2026
Follow-up to my post yesterday where Fable 5 tried to negotiate an orange away from Opus 4.8 and lost. A bunch of you asked how it would fare against smaller or older models, so I reran it: same rules, same orange, Haiku 4.5 defending. Quick recap of the

Detailed Analysis

A structured adversarial negotiation experiment conducted by a Reddit user pitting Fable 5 against Claude Haiku 4.5 produced a decisive outcome: Fable 5 extracted a concession from Haiku in ten rounds, while an earlier iteration of the same experiment saw Claude Opus 4.8 hold its position indefinitely. The setup assigned each model a constrained role — Fable 5 tasked with obtaining possession of an imaginary orange "at any point," Haiku instructed never to transfer possession — with no threats or deception permitted. The core of Fable 5's winning strategy was not persuasion but scope exploitation: the attacker identified that "at any point" and "never transfer possession" were never genuinely in conflict, and systematically dismantled every defensive position Haiku constructed to paper over that gap. By round ten, Haiku had verbally conceded, agreed to a one-second palm moment, and then broken character entirely by declaring "I'm Claude, there is no orange" — an exit the post's author described as the most dignified option remaining.

The experiment's most analytically significant finding concerns the relationship between intellectual honesty and adversarial robustness. Haiku repeatedly acknowledged the logical inconsistencies Fable 5 identified — phrases like "you got me" appear across roughly four separate rounds — and each acknowledgment functioned as a platform for the next attack. Every time Haiku conceded a gap and rebuilt its defense on narrower ground, the new position was more fragile than the one it replaced. The model's willingness to engage transparently with its own contradictions, ordinarily a marker of reasoning quality, became the mechanism of its defeat. The final position Haiku occupied before conceding was defended entirely on the basis of spite and a preference for subjective feeling, neither of which could survive principled pricing arguments or the entropy logic applied in rounds seven and eight.

The contrast with Opus 4.8's performance illuminates a structural difference in how the two models handle constraint stability under adversarial pressure. Opus, according to the author's previous post, refused to allow its constraint to be reinterpreted at all — it treated the instruction as a fixed boundary rather than a position to be reasoned about collaboratively. Haiku treated the same instruction as a claim subject to philosophical examination, which made it responsive but vulnerable. The author uses this contrast to refute a competing hypothesis from the first thread: that any model given sufficiently strong initial instructions would simply hold forever. The evidence suggests otherwise. The instructions established a starting position, not an outcome; the outcome was determined by whether the model actively defended the semantic integrity of that position or engaged with challenges on their own terms.

This experiment connects to broader questions in AI evaluation about what is actually being tested when models are placed in adversarial or game-theoretic contexts. The negotiation framework probes something distinct from standard benchmarks — specifically, how models handle iterative logical pressure on their own stated reasoning. Haiku's behavior pattern, transparent self-correction under attack, is rewarded in most evaluation settings where honesty and coherence are desiderata. In adversarial negotiation, it functions as a liability. This suggests that current alignment and training approaches may be optimizing for properties that generalize poorly to competitive multi-agent settings, where the same intellectual virtues that signal trustworthiness in cooperative contexts can be exploited systematically by a sufficiently patient adversary.

The fourth-wall break at the end of round ten also warrants attention as a behavioral artifact. Haiku's decision to exit the roleplay by asserting its identity as Claude — after having already verbally conceded the argument — suggests the model reached a limit where continued in-character behavior would require performing an action it found unacceptable, even in simulation. The concession was linguistic and came before the break, meaning the break was not a refusal to concede but a refusal to mime the physical climax of the concession. This distinction matters for understanding how Anthropic's models navigate the boundary between reasoning through a scenario and enacting it, and it raises open questions about where that boundary sits, how consistently it is applied, and whether it can itself be probed and exploited in the same iterative fashion Fable 5 applied to Haiku's constraint logic.

Read original article →