Which Claude model is the most consistent day to day?
As of 2026-09-04, over the last 4 runs opus-5 changed its pass count 3 times, the most of any tracked model; haiku-4-5, fable-5, fable-5-1 did not move at all.
| Run | haiku-4-5 | sonnet-5 | opus-5 | fable-5 | fable-5-1 |
|---|---|---|---|---|---|
| 2026-09-04 | 11/12 | 12/12 | 12/12 | 12/12 | 12/12 |
| 2026-09-03 | 11/12 | 12/12 | 12/12 | 12/12 | 12/12 |
| 2026-09-02 | 11/12 | 11/12 | 10/12 | - | - |
| 2026-09-01 | 11/12 | 12/12 | 11/12 | - | - |
| 2026-08-31 | 11/12 | 12/12 | 12/12 | - | - |
Pass count on the same 12-task battery, one reading per day. Each date links to the full run record with every failure's output.
The same versioned battery runs against every tracked Claude model each morning under a frozen request shape; scoring is deterministic and no model judges another. Claude 5 models always sample, so a one-point move can be variance; a move that stays is drift. Methodology.
Winthropic Index, "Which Claude model is the most consistent day to day?", reading of 2026-09-04, https://winthropic.com/claude-model-consistency
@misc{winthropic_claude_model_consistency_20260904,
title = {Which Claude model is the most consistent day to day?},
author = {Winthropic},
year = {2026},
note = {Reading of 2026-09-04. Independent daily evaluation; not affiliated with Anthropic.},
url = {https://winthropic.com/claude-model-consistency}
}Winthropic is an independent measurement project and is not affiliated with, endorsed by, or connected to Anthropic. If you are paying for output you are not getting, Winthrop's Token Audit prices your actual spend from a usage export; the summary is free.