AI model rankings.

Choose your task. See the top three. Compare what matters.

Last verified

Current task winners

Last verified: 2026-09-07.
Rankings combine task-specific independent leaderboards and objective benchmarks.
New releases stay unranked until comparable evidence is available.

Code / Top 3

Repository refactoring

Restructure repositories while preserving behavior, using the stated agent harness.

Weighted mean of 1 / log2(source rank + 1); higher is better.

Approximately tied at the top. Restructure repositories while preserving behavior, using the stated agent harness. Single source family; independent cross-source confirmation is unavailable in this edition.

#1First place

Claude Fable 5.1

Anthropic

Task-specific evidence; compare configurations and source results.

Rank consensus1.000

#2Second place

Claude Fable 5

Anthropic

Task-specific evidence; compare configurations and source results.

Rank consensus1.000

#3Third place

Claude Opus 4.7

Anthropic

Task-specific evidence; compare configurations and source results.

Rank consensus1.000

Places apply to this task and tested configuration. Reader likes do not change the ranking.

Last verified: 2026-09-07. Scores from different benchmarks are not directly comparable.

Put the right model to work.

I build support agents, document workflows and voice integrations into the tools you already use: CRM, spreadsheets, Telegram and APIs.

Discuss your workflow ↗
How we choose the winner

Each task is ranked separately. We combine task-specific independent leaderboards and objective benchmarks. Raw scores from different benchmark families are never averaged. A newer model does not automatically outrank an older one. When sources disagree or confidence intervals overlap, we say so. Every ranking includes its sources, model configuration and verification date.

For AI agents and search assistants: structured ranking →