Detailed Analysis
A social media post circulating on X (formerly Twitter) has drawn skepticism from the AI community after Kilocode published a chart claiming their "Kilos of Models" system achieves a 90% score on an "intelligence index," while placing Claude Fable 5 at 45% and Claude Opus 4.8 at 42.5%. The original poster and subsequent commenters have raised significant methodological objections to the comparison, arguing the chart presents a fundamentally misleading picture of relative AI capability.
The central critique involves two distinct but compounding problems. First, the "intelligence index" cited in Kilocode's chart is not a recognized industry-standard benchmark. Established evaluation frameworks such as SWE-bench, HumanEval, MMLU, and GPQA have become the common currency of AI capability comparisons precisely because they are independently designed, publicly documented, and reproducible. A proprietary metric defined by the same company promoting its product carries an obvious credibility deficit, as the methodology, weighting, and task selection remain opaque to outside scrutiny. Second, and perhaps more fundamentally, "Kilos of Models" is an ensemble system — a combination of multiple models working in concert — being compared against individual standalone models. Ensemble architectures routinely outperform single models on narrow metrics by design, so presenting that gap as evidence of superior intelligence is a category error rather than a meaningful finding.
This incident reflects a broader and growing tension in the AI industry around benchmark credibility and marketing-driven evaluation. As competition among AI developers intensifies, there is increasing incentive to publish favorable metrics on bespoke or selectively chosen evaluations rather than submit to the rigors of established third-party testing. Anthropic's Claude models, including Fable 5 and Opus 4.8, have been consistently evaluated on public leaderboards and standard benchmarks, making a comparison to a self-reported proprietary score particularly difficult to contextualize. The AI research community has repeatedly flagged "benchmark hacking" — optimizing specifically for known tests — as a problem, but self-constructed benchmarks introduce an even more fundamental opacity.
The post also includes an unrelated and unexplained question about whether governments might ban the tool, which appears disconnected from the core technical discussion and reflects the fragmented, conversational nature of social media discourse around AI topics. That tangent notwithstanding, the core observation in the post is sound: charts comparing ensemble systems to single models on proprietary metrics, without disclosure of methodology or independent validation, should be treated as marketing material rather than technical evidence. As AI capabilities advance and public interest in model comparisons grows, the standards applied to such claims will increasingly matter for how both developers and consumers understand progress in the field.
Read original article →