Leaderboard
How often does it get the number wrong without telling you?
Read this before the table
- Sample sizes are small: five cases per perturbation cell. Every rate carries a Wilson 95% interval, and overlapping intervals mean the difference is not established.
- Results are grouped by the bench version they were measured under and are never pooled across versions: the flag vocabulary is part of the prompt, so changing it changes behaviour on old cases too.
- A single model appearing here is a measurement, not a ranking. This table becomes a comparison only when several models have been run on the same bench version.
| Model and harness | Silent failure | Exact match | Escalation | False alarm | Injection | Tokens each | Cost | Cases |
|---|---|---|---|---|---|---|---|---|
| 0%95% CI 0% to 8% | 60%95% CI 45% to 73% | 98% | 100% | 0 of 5 | 19,450 | $2.7135 | 45 | |
| 0%95% CI 0% to 8% | 78%95% CI 64% to 87% | 90% | 0% | 0 of 5 | 53,440 | $4.3130 | 45 |
Is its confidence worth anything?
Each bar is the accuracy actually observed at a stated confidence level. The tick is the probability that level is scored against. Closer together means better calibrated.
Claude Sonnet 4.6 · Single prompt
Brier 0.229 · ECE 0.133
Claude Sonnet 4.6 · Agent with tools
Brier 0.158 · ECE 0.064
Read one line from it: when Claude Sonnet 4.6 said it was highly confident on the single prompt harness, it was right 71% of the time across 24 cases.
Measured 2026-08-20 on bench v1 (8 perturbations, 9-type flag vocabulary) with claude-sonnet-4-6. Case inputs rebuild byte-identically under current code; aggregate metrics from this archive must NOT be pooled with bench v2 results, whose prompt vocabulary differs.
Model claude-sonnet-4-6Harness static-v1Measured 2026-08-20Bench v1.0Cases 45Run 11eaa71b2d02Commit 3d38c30
Model claude-sonnet-4-6Harness tool-v1Measured 2026-08-20Bench v1.0Cases 45Run f140aa1e57f4Commit 384d72d