← Reddit

Same prompt, same model, ten runs: scores from 0.30 to 0.81. This week my skill-testing tool refused to publish its own results, and I shipped the refusal as the report.

Reddit · maverick_man1111 · August 29, 2026
I maintain Driftproof, a small open-source instrument that re-tests agent skills (SKILL.md files) when the model underneath them changes: run the skill's eval suite with and without it, judge each response multiple times, only claim drift when confidence

Detailed Analysis

Detailed analysis coming soon.

Read original article →