Claude Fable 5.1
Anthropic
Task-specific evidence; compare configurations and source results.
Rank consensus1.000
Choose your task. See the top three. Compare what matters.
Last verified
Last verified: 2026-09-07.
Rankings combine task-specific independent leaderboards and objective benchmarks.
New releases stay unranked until comparable evidence is available.
Code / Top 3
Objective coding problems without a terminal agent harness.
Weighted mean of 1 / log2(source rank + 1); higher is better.
Objective coding problems without a terminal agent harness. Single source family; independent cross-source confirmation is unavailable in this edition.
Anthropic
Task-specific evidence; compare configurations and source results.
Rank consensus1.000
Anthropic
Task-specific evidence; compare configurations and source results.
Rank consensus0.631
OpenAI
Task-specific evidence; compare configurations and source results.
Rank consensus0.500
Places apply to this task and tested configuration. Reader likes do not change the ranking.
Last verified: 2026-09-07. Scores from different benchmarks are not directly comparable.
I build support agents, document workflows and voice integrations into the tools you already use: CRM, spreadsheets, Telegram and APIs.
Each task is ranked separately. We combine task-specific independent leaderboards and objective benchmarks. Raw scores from different benchmark families are never averaged. A newer model does not automatically outrank an older one. When sources disagree or confidence intervals overlap, we say so. Every ranking includes its sources, model configuration and verification date.