Disclosure: I built this. Every benchmark I could find measures the model. None of them explain why two people running the same model on the same repo get completely different outcomes — one ships, one burns the session and merges something broken. So I built
Detailed Analysis
Detailed analysis coming soon.
Read original article →