Task-specific AI model rankings — structured for agents and search systems.

Each task has its own top models, evidence, source dates, confidence and disagreement notes. Newer does not automatically mean better.

Last verified: 2026-09-07.
All rankings are derived from task-relevant public or independent evidence.

29 tasks · 39 ranked base models · Methodology 2.0-task-consensus

Consensus methodology

Consensus rankings use only task-relevant sources. Source rankings are normalized within each benchmark before aggregation. Raw scores from different benchmark families are never averaged. Vendor-published benchmarks do not determine rank unless independently reproduced. Models that have not yet been evaluated remain pending rather than being assigned an artificial low score.

method: weighted_reciprocal_rank

formula: weighted_mean(1 / log2(rank + 1))

raw scores averaged: false

source family policy: Within each source family, average benchmark reciprocal ranks; then weight source families. Republished benchmark results are not independent.

configuration policy: One base model per top three. Best reported rank configuration per benchmark contributes once; every raw configuration remains in source_results. Different versions remain separate base models. H3 Max is explicitly a fal post-trained configuration.

admission policy: Final results take priority over preliminary ones. When at least three models have two independent source families, only those enter the consensus top three. Other evaluated models remain candidates; missing evidence is never a bad score.

tie break: Exact consensus ties: coverage, then mean displayed source position, then stable model_id; this order is not statistical superiority.

freshness policy: published_at is the dated leaderboard snapshot, never a model release date or benchmark release version. Unknown publication dates stay null. Evidence older than 30 days is stale; 14 days is a warning.

limitations: Source coverage and confidence describe this captured evidence set. No vendor-reported result determines rank. Unknown API IDs, prices and context sizes are not inferred.

General chat

For broad conversational quality across everyday prompts.

Sources disagree on #1. Human preference is weighted 70%; objective broad ability provides two secondary checks.

Category
text
Metric / benchmark
Text Arena🏆Overall; LiveBench Overall; Intelligence Index
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5
Consensus #2
Claude Fable 5.1
Consensus #3
GPT-6 Astra
Consensus score
0.84463946
Mean rank
3.3333333333333335
Median rank
2
Independent source count
3
Source coverage
1
Confidence
medium
Disagreement status
source_split
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":0.7,"Artificial Analysis":0.15,"LiveBench":0.15}

Top 3

  1. #1 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    0.84463946
    Mean rank
    3.3333333333333335
    Median rank
    2
    Independent source count
    3
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🏆Overall
    Source model name
    claude-fable-5
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 6
    Raw score
    1507
    Unit
    Elo
    Confidence interval
    1502 / 1512
    Sample count
    27189
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    LiveBenchLiveBench Overall
    Source model name
    Claude Fable 5 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    83
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisIntelligence Index
    Source model name
    Claude Fable 5 (with fallback)
    Benchmark version
    4.2
    Source rank
    7
    Rank spread
    Not reported / Not reported
    Raw score
    53
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.65
    Mean rank
    1.6666666666666667
    Median rank
    1
    Independent source count
    3
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🏆Overall
    Source model name
    claude-fable-5.1-max
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 14
    Raw score
    1504
    Unit
    Elo
    Confidence interval
    1493 / 1515
    Sample count
    2906
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    LiveBenchLiveBench Overall
    Source model name
    Claude Fable 5.1 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    83.4
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisIntelligence Index
    Source model name
    Claude Fable 5.1 (max with fallback)
    Benchmark version
    4.2
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    57
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisIntelligence Index
    Source model name
    Claude Fable 5.1 (xhigh with fallback)
    Benchmark version
    4.2
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    56
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Artificial AnalysisIntelligence Index
    Source model name
    Claude Fable 5.1 (high with fallback)
    Benchmark version
    4.2
    Source rank
    4
    Rank spread
    Not reported / Not reported
    Raw score
    54
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Artificial AnalysisIntelligence Index
    Source model name
    Claude Fable 5.1 (medium with fallback)
    Benchmark version
    4.2
    Source rank
    7
    Rank spread
    Not reported / Not reported
    Raw score
    53
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Medium
    Reasoning effort
    medium
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Artificial AnalysisIntelligence Index
    Source model name
    Claude Fable 5.1 (low with fallback)
    Benchmark version
    4.2
    Source rank
    15
    Rank spread
    Not reported / Not reported
    Raw score
    51
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Low
    Reasoning effort
    low
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    false
  3. #3 GPT-6 Astra

    Base model
    GPT-6 Astra
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    3
    Evaluation status
    partially_evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    2
    Source coverage
    0.6666666666666666
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    LiveBenchLiveBench Overall
    Source model name
    GPT-6 Astra Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    82.2
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisIntelligence Index
    Source model name
    GPT-6 Astra (max)
    Benchmark version
    4.2
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    55
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisIntelligence Index
    Source model name
    GPT-6 Astra (xhigh)
    Benchmark version
    4.2
    Source rank
    4
    Rank spread
    Not reported / Not reported
    Raw score
    54
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Artificial AnalysisIntelligence Index
    Source model name
    GPT-6 Astra (high)
    Benchmark version
    4.2
    Source rank
    7
    Rank spread
    Not reported / Not reported
    Raw score
    53
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Artificial AnalysisIntelligence Index
    Source model name
    GPT-6 Astra (medium)
    Benchmark version
    4.2
    Source rank
    12
    Rank spread
    Not reported / Not reported
    Raw score
    52
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Medium
    Reasoning effort
    medium
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Artificial AnalysisIntelligence Index
    Source model name
    GPT-6 Astra (low)
    Benchmark version
    4.2
    Source rank
    22
    Rank spread
    Not reported / Not reported
    Raw score
    49
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Low
    Reasoning effort
    low
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Artificial AnalysisIntelligence Index
    Source model name
    GPT-6 Astra (Non-reasoning)
    Benchmark version
    4.2
    Source rank
    25
    Rank spread
    Not reported / Not reported
    Raw score
    48
    Unit
    index
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:22.585Z
    Evidence freshness
    unknown
    Used in aggregation
    false
Compare this task on the human ranking →

Russian language

For Russian-language conversation and writing.

Approximately tied at the top. Russian-language preference only; overlapping rank ranges limit certainty. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
text
Metric / benchmark
Text Arena🇷🇺Russian
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5
Consensus #2
Muse Spark 1.2
Consensus #3
Claude Fable 5.1
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1}

Top 3

  1. #1 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🇷🇺Russian
    Source model name
    claude-fable-5
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 11
    Raw score
    1521
    Unit
    Elo
    Confidence interval
    1509 / 1533
    Sample count
    2706
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
  2. #2 Muse Spark 1.2

    Base model
    Muse Spark 1.2
    Provider
    Meta
    Official model ID
    Not reported
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🇷🇺Russian
    Source model name
    muse-spark-1.2 (xHigh)
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 39
    Raw score
    1514
    Unit
    Elo
    Confidence interval
    1482 / 1546
    Sample count
    330
    Preliminary
    false
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
  3. #3 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.43067656
    Mean rank
    4
    Median rank
    4
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🇷🇺Russian
    Source model name
    claude-fable-5.1-max
    Benchmark version
    live leaderboard
    Source rank
    4
    Rank spread
    1 / 46
    Raw score
    1510
    Unit
    Elo
    Confidence interval
    1477 / 1543
    Sample count
    348
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
Compare this task on the human ranking →

Expert analysis

For difficult analytical and professional knowledge tasks.

Sources disagree on #1. General expert reasoning; does not establish a winner for individual professions.

Category
text
Metric / benchmark
Text Arena🤓Expert; LiveBench Reasoning; Humanity’s Last Exam
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5
Consensus #2
Claude Fable 5.1
Consensus #3
Claude Opus 4.6
Consensus score
0.73788625
Mean rank
5
Median rank
5
Independent source count
3
Source coverage
0.6666666666666666
Confidence
medium
Disagreement status
source_split
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":0.5,"LiveBench":0.3,"Scale Labs":0.2}

Top 3

  1. #1 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    partially_evaluated
    Consensus score
    0.73788625
    Mean rank
    5
    Median rank
    5
    Independent source count
    2
    Source coverage
    0.6666666666666666
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🤓Expert
    Source model name
    claude-fable-5
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 13
    Raw score
    1548
    Unit
    Elo
    Confidence interval
    1537 / 1559
    Sample count
    2939
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    LiveBenchLiveBench Reasoning
    Source model name
    Claude Fable 5 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    9
    Rank spread
    Not reported / Not reported
    Raw score
    89.7
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.55594559
    Mean rank
    3.3333333333333335
    Median rank
    2
    Independent source count
    3
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🤓Expert
    Source model name
    claude-fable-5.1-max
    Benchmark version
    live leaderboard
    Source rank
    7
    Rank spread
    1 / 65
    Raw score
    1532
    Unit
    Elo
    Confidence interval
    1490 / 1574
    Sample count
    211
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    LiveBenchLiveBench Reasoning
    Source model name
    Claude Fable 5.1 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    91.7
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Scale LabsHumanity’s Last Exam
    Source model name
    Fable 5.1 (xhigh)
    Benchmark version
    current leaderboard
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    46.5
    Unit
    percent
    Confidence interval
    44.5 / 48.5
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Claude Opus 4.6

    Base model
    Claude Opus 4.6
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    High
    Reasoning effort
    high
    Consensus rank
    3
    Evaluation status
    partially_evaluated
    Consensus score
    0.52056426
    Mean rank
    9
    Median rank
    9
    Independent source count
    2
    Source coverage
    0.6666666666666666
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🤓Expert
    Source model name
    claude-opus-4-6-high
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 13
    Raw score
    1547
    Unit
    Elo
    Confidence interval
    1538 / 1556
    Sample count
    6409
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaText Arena🤓Expert
    Source model name
    claude-opus-4-6
    Benchmark version
    live leaderboard
    Source rank
    5
    Rank spread
    1 / 18
    Raw score
    1534
    Unit
    Elo
    Confidence interval
    1526 / 1542
    Sample count
    7644
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    false
    Scale LabsHumanity’s Last Exam
    Source model name
    claude-opus-4-6 (Non-Thinking)
    Benchmark version
    current leaderboard
    Source rank
    16
    Rank spread
    Not reported / Not reported
    Raw score
    19
    Unit
    percent
    Confidence interval
    17.46 / 20.54
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Mathematics

For hard calculations, proofs and multi-step maths.

Sources disagree on #1. Mathematics receives 60%; HLE is a broader reasoning cross-check, not a math-only score.

Category
text
Metric / benchmark
Text Arena🧮Math; LiveBench Mathematics; Humanity’s Last Exam
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5.1
Consensus #2
Claude Fable 5
Consensus #3
Claude Opus 5
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
3
Source coverage
0.6666666666666666
Confidence
medium
Disagreement status
source_split
CI overlap
false
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":0.25,"LiveBench":0.6,"Scale Labs":0.15}

Top 3

  1. #1 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    1
    Evaluation status
    partially_evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    2
    Source coverage
    0.6666666666666666
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    LiveBenchLiveBench Mathematics
    Source model name
    Claude Fable 5.1 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    97
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Scale LabsHumanity’s Last Exam
    Source model name
    Fable 5.1 (xhigh)
    Benchmark version
    current leaderboard
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    46.5
    Unit
    percent
    Confidence interval
    44.5 / 48.5
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    partially_evaluated
    Consensus score
    0.59812463
    Mean rank
    2.5
    Median rank
    2.5
    Independent source count
    2
    Source coverage
    0.6666666666666666
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🧮Math
    Source model name
    claude-fable-5
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 15
    Raw score
    1529
    Unit
    Elo
    Confidence interval
    1513 / 1545
    Sample count
    1331
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    LiveBenchLiveBench Mathematics
    Source model name
    Claude Fable 5 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    4
    Rank spread
    Not reported / Not reported
    Raw score
    96
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Claude Opus 5

    Base model
    Claude Opus 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    3
    Evaluation status
    partially_evaluated
    Consensus score
    0.42086169
    Mean rank
    4.5
    Median rank
    4.5
    Independent source count
    2
    Source coverage
    0.6666666666666666
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena🧮Math
    Source model name
    claude-opus-5-max
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 18
    Raw score
    1529
    Unit
    Elo
    Confidence interval
    1507 / 1551
    Sample count
    740
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaText Arena🧮Math
    Source model name
    claude-opus-5-high
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 16
    Raw score
    1526
    Unit
    Elo
    Confidence interval
    1510 / 1542
    Sample count
    1523
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    false
    LiveBenchLiveBench Mathematics
    Source model name
    Claude 5 Opus Thinking Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    7
    Rank spread
    Not reported / Not reported
    Raw score
    95.7
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Instruction following

For prompts with detailed constraints and required formats.

Sources disagree on #1. Preference, objective instruction following and IFBench measure different constraints; see each result.

Category
text
Metric / benchmark
Text Arena📝Instruction Following; LiveBench Instruction Following; IFBench
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Grok 4.3
Consensus #2
Claude Fable 5
Consensus #3
Gemini 3.1 Pro Preview
Consensus score
0.44489779
Mean rank
49.333333333333336
Median rank
40
Independent source count
3
Source coverage
1
Confidence
medium
Disagreement status
source_split
CI overlap
false
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1,"Artificial Analysis":1,"LiveBench":1}

Top 3

  1. #1 Grok 4.3

    Base model
    Grok 4.3
    Provider
    SpaceXAI
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    0.44489779
    Mean rank
    49.333333333333336
    Median rank
    40
    Independent source count
    3
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena📝Instruction Following
    Source model name
    grok-4.3
    Benchmark version
    live leaderboard
    Source rank
    107
    Rank spread
    86 / 130
    Raw score
    1416
    Unit
    Elo
    Confidence interval
    1411 / 1421
    Sample count
    22801
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    LiveBenchLiveBench Instruction Following
    Source model name
    Grok 4.3
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    40
    Rank spread
    Not reported / Not reported
    Raw score
    62.8
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisIFBench
    Source model name
    Grok 4.3 (medium)
    Benchmark version
    58 constraints; visible top-10 chart
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    83.3
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Medium
    Reasoning effort
    medium
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisIFBench
    Source model name
    Grok 4.3 (high)
    Benchmark version
    58 constraints; visible top-10 chart
    Source rank
    5
    Rank spread
    Not reported / Not reported
    Raw score
    81.3
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    false
  2. #2 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.42540059
    Mean rank
    6
    Median rank
    6
    Independent source count
    3
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena📝Instruction Following
    Source model name
    claude-fable-5
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 7
    Raw score
    1512
    Unit
    Elo
    Confidence interval
    1505 / 1519
    Sample count
    9607
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    LiveBenchLiveBench Instruction Following
    Source model name
    Claude Fable 5 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    6
    Rank spread
    Not reported / Not reported
    Raw score
    75.8
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisIFBench
    Source model name
    Claude Fable 5 (with fallback)
    Benchmark version
    58 constraints; visible top-10 chart
    Source rank
    10
    Rank spread
    Not reported / Not reported
    Raw score
    63.5
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Gemini 3.1 Pro Preview

    Base model
    Gemini 3.1 Pro Preview
    Provider
    Google
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    partially_evaluated
    Consensus score
    0.375
    Mean rank
    9
    Median rank
    9
    Independent source count
    2
    Source coverage
    0.6666666666666666
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText Arena📝Instruction Following
    Source model name
    gemini-3.1-pro-preview
    Benchmark version
    live leaderboard
    Source rank
    15
    Rank spread
    9 / 31
    Raw score
    1481
    Unit
    Elo
    Confidence interval
    1476 / 1486
    Sample count
    34766
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    LiveBenchLiveBench Instruction Following
    Source model name
    Gemini 3.1 Pro Preview High
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    79.1
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Copywriting

For persuasive, creative and polished marketing text.

Sources disagree on #1. Creative writing and writing/literature/language without style control. These are proxies for copywriting, not conversion tests. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
text
Metric / benchmark
Text Arena✍️Creative Writing; Text Arena✍️Writing, Literature, & Language
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5
Consensus #2
Claude Fable 5.1
Consensus #3
Claude Opus 4.6
Consensus score
0.81546488
Mean rank
1.5
Median rank
1.5
Independent source count
1
Source coverage
1
Confidence
low
Disagreement status
source_split
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1}

Top 3

  1. #1 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    0.81546488
    Mean rank
    1.5
    Median rank
    1.5
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    low
    Source-specific results
    ArenaText Arena✍️Creative Writing
    Source model name
    claude-fable-5
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 6
    Raw score
    1504
    Unit
    Elo
    Confidence interval
    1495 / 1513
    Sample count
    5480
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaText Arena✍️Writing, Literature, & Language
    Source model name
    claude-fable-5
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 8
    Raw score
    1505
    Unit
    Elo
    Confidence interval
    1497 / 1513
    Sample count
    7328
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
  2. #2 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.67810359
    Mean rank
    3.5
    Median rank
    3.5
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    low
    Source-specific results
    ArenaText Arena✍️Creative Writing
    Source model name
    claude-fable-5.1-max
    Benchmark version
    live leaderboard
    Source rank
    6
    Rank spread
    1 / 30
    Raw score
    1487
    Unit
    Elo
    Confidence interval
    1462 / 1512
    Sample count
    608
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaText Arena✍️Writing, Literature, & Language
    Source model name
    claude-fable-5.1-max
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 7
    Raw score
    1526
    Unit
    Elo
    Confidence interval
    1504 / 1548
    Sample count
    782
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
  3. #3 Claude Opus 4.6

    Base model
    Claude Opus 4.6
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    High
    Reasoning effort
    high
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.56546488
    Mean rank
    2.5
    Median rank
    2.5
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    low
    Source-specific results
    ArenaText Arena✍️Creative Writing
    Source model name
    claude-opus-4-6-high
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 6
    Raw score
    1500
    Unit
    Elo
    Confidence interval
    1493 / 1507
    Sample count
    12678
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaText Arena✍️Creative Writing
    Source model name
    claude-opus-4-6
    Benchmark version
    live leaderboard
    Source rank
    9
    Rank spread
    3 / 22
    Raw score
    1479
    Unit
    Elo
    Confidence interval
    1472 / 1486
    Sample count
    12724
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    false
    ArenaText Arena✍️Writing, Literature, & Language
    Source model name
    claude-opus-4-6-high
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 9
    Raw score
    1502
    Unit
    Elo
    Confidence interval
    1496 / 1508
    Sample count
    18186
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaText Arena✍️Writing, Literature, & Language
    Source model name
    claude-opus-4-6
    Benchmark version
    live leaderboard
    Source rank
    7
    Rank spread
    2 / 10
    Raw score
    1494
    Unit
    Elo
    Confidence interval
    1488 / 1500
    Sample count
    18365
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    false
Compare this task on the human ranking →

Search and research

For web search, evidence gathering and sourced answers.

Approximately tied at the top. Search-enabled configurations only. General intelligence is not search evidence. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
text
Metric / benchmark
Search Arena
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
GPT-5.6 Sol
Consensus #2
Claude Opus 4.6
Consensus #3
GPT-5.5
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1}

Top 3

  1. #1 GPT-5.6 Sol

    Base model
    GPT-5.6 Sol
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    ArenaSearch Arena
    Source model name
    gpt-5.6-sol-xhigh
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 2
    Raw score
    1257
    Unit
    Elo
    Confidence interval
    1250 / 1264
    Sample count
    29663
    Preliminary
    false
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Source date
    2026-08-24
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    warning
    Used in aggregation
    true
  2. #2 Claude Opus 4.6

    Base model
    Claude Opus 4.6
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    ArenaSearch Arena
    Source model name
    claude-opus-4-6-search
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 2
    Raw score
    1253
    Unit
    Elo
    Confidence interval
    1248 / 1258
    Sample count
    134699
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-08-24
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    warning
    Used in aggregation
    true
  3. #3 GPT-5.5

    Base model
    GPT-5.5
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    ArenaSearch Arena
    Source model name
    gpt-5.5-search
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    3 / 5
    Raw score
    1242
    Unit
    Elo
    Confidence interval
    1237 / 1247
    Sample count
    89873
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-08-24
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    warning
    Used in aggregation
    true

Pending evaluation

  • GPT-6 Astra: pending_evaluation. New release — awaiting comparable task-level evidence.
  • Claude Fable 5.1: pending_evaluation. New release — awaiting comparable task-level evidence.
Compare this task on the human ranking →

PDF and documents

For understanding long documents, PDFs and mixed document content.

Approximately tied at the top. Document Arena preference; newly released models without a document evaluation remain pending. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
text
Metric / benchmark
Document Arena
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Opus 5
Consensus #2
Claude Opus 4.6
Consensus #3
Claude Fable 5
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1}

Top 3

  1. #1 Claude Opus 5

    Base model
    Claude Opus 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    High
    Reasoning effort
    high
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    ArenaDocument Arena
    Source model name
    claude-opus-5-high
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 4
    Raw score
    1520
    Unit
    Elo
    Confidence interval
    1505 / 1535
    Sample count
    1663
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Claude Opus 4.6

    Base model
    Claude Opus 4.6
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    ArenaDocument Arena
    Source model name
    claude-opus-4-6
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 6
    Raw score
    1510
    Unit
    Elo
    Confidence interval
    1504 / 1516
    Sample count
    37271
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
    ArenaDocument Arena
    Source model name
    claude-opus-4-6-high
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 6
    Raw score
    1506
    Unit
    Elo
    Confidence interval
    1499 / 1513
    Sample count
    24973
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    false
  3. #3 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.43067656
    Mean rank
    4
    Median rank
    4
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    ArenaDocument Arena
    Source model name
    claude-fable-5
    Benchmark version
    live leaderboard
    Source rank
    4
    Rank spread
    1 / 7
    Raw score
    1504
    Unit
    Elo
    Confidence interval
    1495 / 1513
    Sample count
    5792
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true

Pending evaluation

  • GPT-6 Astra: pending_evaluation. New release — awaiting comparable task-level evidence.
  • Claude Fable 5.1: pending_evaluation. New release — awaiting comparable task-level evidence.
Compare this task on the human ranking →

Terminal software engineering

For repository work, terminal coding and software engineering.

Approximately tied at the top. Terminal-Bench 4.0 with the stated agent harness; does not measure standalone code reasoning. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
code
Metric / benchmark
Terminal-Bench
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
GPT-6 Astra
Consensus #2
Claude Fable 5.1
Consensus #3
Claude Opus 5
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Terminal-Bench":1}

Top 3

  1. #1 GPT-6 Astra

    Base model
    GPT-6 Astra
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Codex / (max)
    Reasoning effort
    max
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Terminal-BenchTerminal-Bench
    Source model name
    GPT-6 Astra (max)
    Benchmark version
    4.0
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    58.2
    Unit
    percent
    Confidence interval
    55.400000000000006 / 61
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Codex / (max)
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.178Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Claude Code / (max)
    Reasoning effort
    max
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Terminal-BenchTerminal-Bench
    Source model name
    Fable 5.1 (max)
    Benchmark version
    4.0
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    57.9
    Unit
    percent
    Confidence interval
    54.1 / 61.699999999999996
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Claude Code / (max)
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.178Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Claude Opus 5

    Base model
    Claude Opus 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Claude Code / (max)
    Reasoning effort
    max
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Terminal-BenchTerminal-Bench
    Source model name
    Opus 5 (max)
    Benchmark version
    4.0
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    51.8
    Unit
    percent
    Confidence interval
    48.4 / 55.199999999999996
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Claude Code / (max)
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.178Z
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Web development

For building and editing interactive web interfaces.

Approximately tied at the top. WebDev overall preference. The leading models have overlapping rank spreads. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
code
Metric / benchmark
Code Arena | WebDev🏆Overall
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
GPT-6 Astra
Consensus #2
Claude Fable 5.1
Consensus #3
Claude Opus 5
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1}

Top 3

  1. #1 GPT-6 Astra

    Base model
    GPT-6 Astra
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaCode Arena | WebDev🏆Overall
    Source model name
    gpt-6-astra-max
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 2
    Raw score
    1797
    Unit
    Elo
    Confidence interval
    1773 / 1821
    Sample count
    1199
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
  2. #2 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaCode Arena | WebDev🏆Overall
    Source model name
    claude-fable-5.1-max
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 2
    Raw score
    1762
    Unit
    Elo
    Confidence interval
    1746 / 1778
    Sample count
    2275
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
  3. #3 Claude Opus 5

    Base model
    Claude Opus 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaCode Arena | WebDev🏆Overall
    Source model name
    claude-opus-5-max
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    3 / 6
    Raw score
    1688
    Unit
    Elo
    Confidence interval
    1680 / 1696
    Sample count
    10904
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaCode Arena | WebDev🏆Overall
    Source model name
    claude-opus-5-high
    Benchmark version
    live leaderboard
    Source rank
    7
    Rank spread
    5 / 7
    Raw score
    1661
    Unit
    Elo
    Confidence interval
    1654 / 1668
    Sample count
    10928
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    false
Compare this task on the human ranking →

AI agents

For autonomous multi-step work with tools and external environments.

Approximately tied at the top. Agent Arena overall plus targeted clarification ability; no claim about every agent workload.

Category
code
Metric / benchmark
Agent Arena🏆Overall; HiL-Bench
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5.1
Consensus #2
Claude Opus 5
Consensus #3
Claude Fable 5
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
2
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":0.75,"Scale Labs":0.25}

Top 3

  1. #1 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaAgent Arena🏆Overall
    Source model name
    Claude Fable 5.1 (Max)
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 3
    Raw score
    15.87
    Unit
    percent
    Confidence interval
    13.03 / 18.71
    Sample count
    6796
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Scale LabsHiL-Bench
    Source model name
    Claude Fable 5.1
    Benchmark version
    current leaderboard
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    61.5
    Unit
    percent
    Confidence interval
    55.03 / 67.97
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Claude Opus 5

    Base model
    Claude Opus 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    High
    Reasoning effort
    high
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.72319732
    Mean rank
    1.5
    Median rank
    1.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaAgent Arena🏆Overall
    Source model name
    Claude Opus 5 (High)
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 4
    Raw score
    12.68
    Unit
    percent
    Confidence interval
    10.95 / 14.41
    Sample count
    22594
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaAgent Arena🏆Overall
    Source model name
    Claude Opus 5 (Max)
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 6
    Raw score
    11.67
    Unit
    percent
    Confidence interval
    9.71 / 13.629999999999999
    Sample count
    18112
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    false
    Scale LabsHiL-Bench
    Source model name
    Claude Opus 5
    Benchmark version
    current leaderboard
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    57
    Unit
    percent
    Confidence interval
    51.519999999999996 / 62.480000000000004
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    High
    Reasoning effort
    high
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.57300742
    Mean rank
    2.5
    Median rank
    2.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaAgent Arena🏆Overall
    Source model name
    Claude Fable 5 (High)
    Benchmark version
    live leaderboard
    Source rank
    4
    Rank spread
    2 / 7
    Raw score
    10.23
    Unit
    percent
    Confidence interval
    8.77 / 11.690000000000001
    Sample count
    36204
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Scale LabsHiL-Bench
    Source model name
    Claude Fable 5
    Benchmark version
    current leaderboard
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    56.33
    Unit
    percent
    Confidence interval
    50.83 / 61.83
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Image understanding

For reasoning about photos, diagrams and visual content.

Approximately tied at the top. Visual question-answering preference; does not establish performance on every vision subtask. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
image
Metric / benchmark
Vision Arena🏆Overall
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5
Consensus #2
Claude Opus 4.7
Consensus #3
Qwen3.8 Max
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1}

Top 3

  1. #1 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaVision Arena🏆Overall
    Source model name
    claude-fable-5
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 6
    Raw score
    1313
    Unit
    Elo
    Confidence interval
    1305 / 1321
    Sample count
    10002
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-08-27
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
  2. #2 Claude Opus 4.7

    Base model
    Claude Opus 4.7
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    High
    Reasoning effort
    high
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaVision Arena🏆Overall
    Source model name
    claude-opus-4-7-high
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 11
    Raw score
    1301
    Unit
    Elo
    Confidence interval
    1294 / 1308
    Sample count
    21137
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    2026-08-27
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaVision Arena🏆Overall
    Source model name
    claude-opus-4-7
    Benchmark version
    live leaderboard
    Source rank
    4
    Rank spread
    1 / 14
    Raw score
    1299
    Unit
    Elo
    Confidence interval
    1292 / 1306
    Sample count
    21457
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-08-27
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    false
  3. #3 Qwen3.8 Max

    Base model
    Qwen3.8 Max
    Provider
    Alibaba
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaVision Arena🏆Overall
    Source model name
    qwen3.8-max
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 15
    Raw score
    1300
    Unit
    Elo
    Confidence interval
    1292 / 1308
    Sample count
    7244
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-08-27
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
Compare this task on the human ranking →

Text from image (OCR)

For reading text embedded in images and scanned content.

Approximately tied at the top. OCR task slice; includes confidence intervals and vote counts. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
image
Metric / benchmark
Vision — OCR
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5
Consensus #2
Qwen3.8 Max
Consensus #3
Claude Opus 4.7
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1}

Top 3

  1. #1 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaVision — OCR
    Source model name
    claude-fable-5
    Benchmark version
    Not reported
    Source rank
    1
    Rank spread
    1 / 5
    Raw score
    1329
    Unit
    Elo
    Confidence interval
    1320 / 1338
    Sample count
    7174
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:10:38.778Z
    Evidence freshness
    fresh
    Used in aggregation
    true
  2. #2 Qwen3.8 Max

    Base model
    Qwen3.8 Max
    Provider
    Alibaba
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaVision — OCR
    Source model name
    qwen3.8-max
    Benchmark version
    Not reported
    Source rank
    2
    Rank spread
    1 / 13
    Raw score
    1315
    Unit
    Elo
    Confidence interval
    1306 / 1324
    Sample count
    5092
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:10:38.778Z
    Evidence freshness
    fresh
    Used in aggregation
    true
  3. #3 Claude Opus 4.7

    Base model
    Claude Opus 4.7
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaVision — OCR
    Source model name
    claude-opus-4-7
    Benchmark version
    Not reported
    Source rank
    3
    Rank spread
    1 / 12
    Raw score
    1314
    Unit
    Elo
    Confidence interval
    1307 / 1321
    Sample count
    15512
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:10:38.778Z
    Evidence freshness
    fresh
    Used in aggregation
    true
    ArenaVision — OCR
    Source model name
    claude-opus-4-7-high
    Benchmark version
    Not reported
    Source rank
    4
    Rank spread
    1 / 12
    Raw score
    1314
    Unit
    Elo
    Confidence interval
    1307 / 1321
    Sample count
    15236
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:10:38.778Z
    Evidence freshness
    fresh
    Used in aggregation
    false
Compare this task on the human ranking →

Image generation

For generating images from natural-language prompts.

Independent image preference sources; configurations are retained separately.

Category
image
Metric / benchmark
Text-to-Image Arena🏆Overall; Text to Image Leaderboard - Top AI Image Models | Artificial Analysis
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
GPT Image 2
Consensus #2
MAI-Image 2.6
Consensus #3
Reve 2.1
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
2
Source coverage
1
Confidence
medium
Disagreement status
clear_winner
CI overlap
false
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1,"Artificial Analysis":1}

Top 3

  1. #1 GPT Image 2

    Base model
    GPT Image 2
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Medium
    Reasoning effort
    medium
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText-to-Image Arena🏆Overall
    Source model name
    gpt-image-2 (medium)
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 1
    Raw score
    1382
    Unit
    Elo
    Confidence interval
    1378 / 1386
    Sample count
    77830
    Preliminary
    false
    Variant
    Medium
    Reasoning effort
    medium
    Source date
    2026-09-04
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisText to Image Leaderboard - Top AI Image Models | Artificial Analysis
    Source model name
    GPT Image 2 (high)
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 1
    Raw score
    1178
    Unit
    Elo
    Confidence interval
    1168 / 1188
    Sample count
    14585
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 MAI-Image 2.6

    Base model
    MAI-Image 2.6
    Provider
    Microsoft AI
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText-to-Image Arena🏆Overall
    Source model name
    mai-image-2.6
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    2 / 3
    Raw score
    1332
    Unit
    Elo
    Confidence interval
    1325 / 1339
    Sample count
    10086
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-04
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisText to Image Leaderboard - Top AI Image Models | Artificial Analysis
    Source model name
    MAI-Image-2.6
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    2 / 2
    Raw score
    1149
    Unit
    Elo
    Confidence interval
    1137 / 1161
    Sample count
    6351
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Reve 2.1

    Base model
    Reve 2.1
    Provider
    Reve
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.46533828
    Mean rank
    3.5
    Median rank
    3.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText-to-Image Arena🏆Overall
    Source model name
    reve-2.1
    Benchmark version
    live leaderboard
    Source rank
    4
    Rank spread
    3 / 4
    Raw score
    1301
    Unit
    Elo
    Confidence interval
    1293 / 1309
    Sample count
    7576
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-04
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisText to Image Leaderboard - Top AI Image Models | Artificial Analysis
    Source model name
    Reve 2.1
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    3 / 4
    Raw score
    1127
    Unit
    Elo
    Confidence interval
    1118 / 1136
    Sample count
    16045
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Image editing

For editing existing images from text instructions.

Sources disagree on #1. Both GPT Image 2 and MAI-Image 2.6 are top tier.

Category
image
Metric / benchmark
Image Edit Arena🏆Single Image Edit; Image Editing Leaderboard - Top AI Image Models | Artificial Analysis
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
GPT Image 2
Consensus #2
MAI-Image 2.6
Consensus #3
Muse Image
Consensus score
0.81546488
Mean rank
1.5
Median rank
1.5
Independent source count
2
Source coverage
1
Confidence
medium
Disagreement status
source_split
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1,"Artificial Analysis":1}

Top 3

  1. #1 GPT Image 2

    Base model
    GPT Image 2
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Medium
    Reasoning effort
    medium
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    0.81546488
    Mean rank
    1.5
    Median rank
    1.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaImage Edit Arena🏆Single Image Edit
    Source model name
    gpt-image-2 (medium)
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 1
    Raw score
    1461
    Unit
    Elo
    Confidence interval
    1457 / 1465
    Sample count
    228827
    Preliminary
    false
    Variant
    Medium
    Reasoning effort
    medium
    Source date
    2026-09-03
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisImage Editing Leaderboard - Top AI Image Models | Artificial Analysis
    Source model name
    GPT Image 2 (high)
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 2
    Raw score
    1117
    Unit
    Elo
    Confidence interval
    1107 / 1127
    Sample count
    11046
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 MAI-Image 2.6

    Base model
    MAI-Image 2.6
    Provider
    Microsoft AI
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.81546488
    Mean rank
    1.5
    Median rank
    1.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaImage Edit Arena🏆Single Image Edit
    Source model name
    mai-image-2.6
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    2 / 3
    Raw score
    1439
    Unit
    Elo
    Confidence interval
    1429 / 1449
    Sample count
    5086
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-03
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisImage Editing Leaderboard - Top AI Image Models | Artificial Analysis
    Source model name
    MAI-Image-2.6
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 2
    Raw score
    1122
    Unit
    Elo
    Confidence interval
    1112 / 1132
    Sample count
    11173
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Muse Image

    Base model
    Muse Image
    Provider
    Meta
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.46533828
    Mean rank
    3.5
    Median rank
    3.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaImage Edit Arena🏆Single Image Edit
    Source model name
    muse-image
    Benchmark version
    live leaderboard
    Source rank
    4
    Rank spread
    4 / 5
    Raw score
    1403
    Unit
    Elo
    Confidence interval
    1398 / 1408
    Sample count
    76663
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-03
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisImage Editing Leaderboard - Top AI Image Models | Artificial Analysis
    Source model name
    Muse Image
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    3 / 6
    Raw score
    1105
    Unit
    Elo
    Confidence interval
    1095 / 1115
    Sample count
    5493
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Text → video

For generating complete videos from text prompts.

Sources disagree on #1. Independent preference sources disagree; overlapping intervals prevent a definitive winner.

Category
video
Metric / benchmark
Text-to-Video Arena; Text to Video Leaderboard - Top AI Video Models | Artificial Analysis
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Wan 3.0
Consensus #2
Gemini Omni Flash
Consensus #3
MiniMax H3
Consensus score
0.75
Mean rank
2
Median rank
2
Independent source count
2
Source coverage
1
Confidence
medium
Disagreement status
source_split
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1,"Artificial Analysis":1}

Top 3

  1. #1 Wan 3.0

    Base model
    Wan 3.0
    Provider
    Alibaba
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    0.75
    Mean rank
    2
    Median rank
    2
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText-to-Video Arena
    Source model name
    wan3.0
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 7
    Raw score
    1494
    Unit
    Elo
    Confidence interval
    1475 / 1513
    Sample count
    1167
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-04
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisText to Video Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    Wan 3.0
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 3
    Raw score
    1239
    Unit
    Elo
    Confidence interval
    1229 / 1249
    Sample count
    5591
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Gemini Omni Flash

    Base model
    Gemini Omni Flash
    Provider
    Google
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText-to-Video Arena
    Source model name
    gemini-omni-flash
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 5
    Raw score
    1511
    Unit
    Elo
    Confidence interval
    1501 / 1521
    Sample count
    21934
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-04
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisText to Video Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    Gemini Omni Flash
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 3
    Raw score
    1238
    Unit
    Elo
    Confidence interval
    1232 / 1244
    Sample count
    18134
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 MiniMax H3

    Base model
    MiniMax H3
    Provider
    MiniMax
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.40773244
    Mean rank
    5.5
    Median rank
    5.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaText-to-Video Arena
    Source model name
    minimax-h3
    Benchmark version
    live leaderboard
    Source rank
    8
    Rank spread
    6 / 9
    Raw score
    1462
    Unit
    Elo
    Confidence interval
    1452 / 1472
    Sample count
    7648
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-04
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisText to Video Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    Minimax H3 Max (post-trained by fal)
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 4
    Raw score
    1235
    Unit
    Elo
    Confidence interval
    1225 / 1245
    Sample count
    5451
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisText to Video Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    MiniMax H3
    Benchmark version
    live leaderboard
    Source rank
    4
    Rank spread
    4 / 5
    Raw score
    1227
    Unit
    Elo
    Confidence interval
    1220 / 1234
    Sample count
    8993
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    false
Compare this task on the human ranking →

Image → video

For animating a still image into a generated video.

Sources disagree on ranking order. H3 Max is a fal post-trained configuration. Source-specific configuration evidence is not interchangeable.

Category
video
Metric / benchmark
Image-to-Video Arena; Image to Video Leaderboard - Top AI Video Models | Artificial Analysis
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
MiniMax H3
Consensus #2
Dreamina Seedance 2.0
Consensus #3
Wan 3.0
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
2
Source coverage
1
Confidence
medium
Disagreement status
source_split
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1,"Artificial Analysis":1}

Top 3

  1. #1 MiniMax H3

    Base model
    MiniMax H3
    Provider
    MiniMax
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaImage-to-Video Arena
    Source model name
    minimax-h3
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 3
    Raw score
    1497
    Unit
    Elo
    Confidence interval
    1491 / 1503
    Sample count
    36137
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisImage to Video Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    Minimax H3 Max (post-trained by fal)
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 2
    Raw score
    1201
    Unit
    Elo
    Confidence interval
    1192 / 1210
    Sample count
    5587
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
    Artificial AnalysisImage to Video Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    MiniMax H3
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    2 / 4
    Raw score
    1187
    Unit
    Elo
    Confidence interval
    1179 / 1195
    Sample count
    7465
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    false
  2. #2 Dreamina Seedance 2.0

    Base model
    Dreamina Seedance 2.0
    Provider
    Bytedance
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.50889128
    Mean rank
    3.5
    Median rank
    3.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaImage-to-Video Arena
    Source model name
    dreamina-seedance-2.0-720p
    Benchmark version
    live leaderboard
    Source rank
    5
    Rank spread
    2 / 5
    Raw score
    1477
    Unit
    Elo
    Confidence interval
    1469 / 1485
    Sample count
    113611
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisImage to Video Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    Dreamina Seedance 2.0 720p
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    2 / 3
    Raw score
    1192
    Unit
    Elo
    Confidence interval
    1186 / 1198
    Sample count
    18094
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Wan 3.0

    Base model
    Wan 3.0
    Provider
    Alibaba
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.4434264
    Mean rank
    4
    Median rank
    4
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaImage-to-Video Arena
    Source model name
    wan3.0
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 6
    Raw score
    1481
    Unit
    Elo
    Confidence interval
    1465 / 1497
    Sample count
    1401
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-09-02
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisImage to Video Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    Wan 3.0
    Benchmark version
    live leaderboard
    Source rank
    5
    Rank spread
    4 / 5
    Raw score
    1176
    Unit
    Elo
    Confidence interval
    1168 / 1184
    Sample count
    7791
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Video editing

For instruction-driven edits and transformations of existing video.

Approximately tied at the top. Both sources put Wan 3.0 first, with uncertainty on Arena.

Category
video
Metric / benchmark
Video Edit Arena; Video Editing Leaderboard - Top AI Video Models | Artificial Analysis
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Wan 3.0
Consensus #2
MiniMax H3
Consensus #3
Gemini Omni Flash
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
2
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1,"Artificial Analysis":1}

Top 3

  1. #1 Wan 3.0

    Base model
    Wan 3.0
    Provider
    Alibaba
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaVideo Edit Arena
    Source model name
    wan3.0
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 3
    Raw score
    1414
    Unit
    Elo
    Confidence interval
    1388 / 1440
    Sample count
    463
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-08-26
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisVideo Editing Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    Wan 3.0
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 1
    Raw score
    1190
    Unit
    Elo
    Confidence interval
    1183 / 1197
    Sample count
    5315
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 MiniMax H3

    Base model
    MiniMax H3
    Provider
    MiniMax
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.56546488
    Mean rank
    2.5
    Median rank
    2.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaVideo Edit Arena
    Source model name
    minimax-h3
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    1 / 5
    Raw score
    1392
    Unit
    Elo
    Confidence interval
    1373 / 1411
    Sample count
    962
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-08-26
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisVideo Editing Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    MiniMax H3
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    2 / 3
    Raw score
    1129
    Unit
    Elo
    Confidence interval
    1124 / 1134
    Sample count
    13014
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Gemini Omni Flash

    Base model
    Gemini Omni Flash
    Provider
    Google
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.46533828
    Mean rank
    3.5
    Median rank
    3.5
    Independent source count
    2
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaVideo Edit Arena
    Source model name
    gemini-omni-flash
    Benchmark version
    live leaderboard
    Source rank
    4
    Rank spread
    3 / 5
    Raw score
    1367
    Unit
    Elo
    Confidence interval
    1352 / 1382
    Sample count
    2109
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    2026-08-26
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
    Artificial AnalysisVideo Editing Leaderboard - Top AI Video Models | Artificial Analysis
    Source model name
    Gemini Omni Flash
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    2 / 3
    Raw score
    1125
    Unit
    Elo
    Confidence interval
    1120 / 1130
    Sample count
    15995
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Speech to text

For transcribing spoken audio into text.

AA-WER transcription accuracy only. Lower WER is better; latency is a separate criterion. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
audio
Metric / benchmark
AA-WER Non-streaming
Consensus metric direction
higher_is_better
Raw metric direction
lower_is_better
Consensus #1
Fun-Realtime-ASR-preview
Consensus #2
MAI-Transcribe-2
Consensus #3
ElevenLabs Scribe v2
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
clear_winner
CI overlap
false
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Artificial Analysis":1}

Top 3

  1. #1 Fun-Realtime-ASR-preview

    Base model
    Fun-Realtime-ASR-preview
    Provider
    Alibaba
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisAA-WER Non-streaming
    Source model name
    Fun-Realtime-ASR-preview
    Benchmark version
    Not reported
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    1.7
    Unit
    percent WER
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 MAI-Transcribe-2

    Base model
    MAI-Transcribe-2
    Provider
    Microsoft AI
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisAA-WER Non-streaming
    Source model name
    MAI-Transcribe-2
    Benchmark version
    Not reported
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    2
    Unit
    percent WER
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 ElevenLabs Scribe v2

    Base model
    ElevenLabs Scribe v2
    Provider
    ElevenLabs
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisAA-WER Non-streaming
    Source model name
    ElevenLabs Scribe v2
    Benchmark version
    Not reported
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    2.2
    Unit
    percent WER
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Text to speech

For generating natural spoken audio from text.

Approximately tied at the top. Provider Voice Arena measures preference with each provider’s selected voices. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
audio
Metric / benchmark
Text to Speech Leaderboard - Top AI Speech Models | Artificial Analysis
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Sonic 3.6
Consensus #2
Realtime TTS-2
Consensus #3
Qwen-Audio-3.0-TTS-Plus
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Artificial Analysis":1}

Top 3

  1. #1 Sonic 3.6

    Base model
    Sonic 3.6
    Provider
    Cartesia
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisText to Speech Leaderboard - Top AI Speech Models | Artificial Analysis
    Source model name
    Sonic 3.6
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 1
    Raw score
    1282
    Unit
    Elo
    Confidence interval
    1265 / 1299
    Sample count
    1790
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Realtime TTS-2

    Base model
    Realtime TTS-2
    Provider
    Inworld
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisText to Speech Leaderboard - Top AI Speech Models | Artificial Analysis
    Source model name
    Realtime TTS-2
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    2 / 4
    Raw score
    1252
    Unit
    Elo
    Confidence interval
    1234 / 1270
    Sample count
    1094
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Qwen-Audio-3.0-TTS-Plus

    Base model
    Qwen-Audio-3.0-TTS-Plus
    Provider
    Alibaba
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisText to Speech Leaderboard - Top AI Speech Models | Artificial Analysis
    Source model name
    Qwen-Audio-3.0-TTS-Plus
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    2 / 5
    Raw score
    1241
    Unit
    Elo
    Confidence interval
    1227 / 1255
    Sample count
    2398
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Voice agents

For voice agents completing customer-service tasks through spoken interaction.

Voice agents ranked by customer-service task completion (tau-Voice); conversation is a separate task. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
audio
Metric / benchmark
Voice Agentic Performance
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Grok Voice Think Fast 2.0
Consensus #2
Qwen Audio 3.0 Realtime Plus
Consensus #3
Grok Voice Think Fast 1.0
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
clear_winner
CI overlap
false
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Artificial Analysis":1}

Top 3

  1. #1 Grok Voice Think Fast 2.0

    Base model
    Grok Voice Think Fast 2.0
    Provider
    SpaceXAI
    Official model ID
    Not reported
    Variant
    High
    Reasoning effort
    high
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisVoice Agentic Performance
    Source model name
    Grok Voice Think Fast 2.0 High
    Benchmark version
    current components
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    56.5
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Speech to Speech Index %
    79.0%
    Speech Reasoning %
    97%
    Conversational Dynamics %
    95.1%
    Agentic Performance %
    56.5%
    Arena Preference Elo
    908
    Task Success Rate %
    94.7%
    Time to First Audio Seconds
    0.70
  2. #2 Qwen Audio 3.0 Realtime Plus

    Base model
    Qwen Audio 3.0 Realtime Plus
    Provider
    Alibaba Cloud
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisVoice Agentic Performance
    Source model name
    Qwen Audio 3.0 Realtime Plus
    Benchmark version
    current components
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    54.6
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Speech to Speech Index %
    66.8%
    Speech Reasoning %
    99%
    Conversational Dynamics %
    98.4%
    Agentic Performance %
    54.6%
    Arena Preference Elo
    699
    Task Success Rate %
    77.8%
    Time to First Audio Seconds
    1.54
  3. #3 Grok Voice Think Fast 1.0

    Base model
    Grok Voice Think Fast 1.0
    Provider
    SpaceXAI
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisVoice Agentic Performance
    Source model name
    Grok Voice Think Fast 1.0
    Benchmark version
    current components
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    52.1
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Speech to Speech Index %
    72.3%
    Speech Reasoning %
    97%
    Conversational Dynamics %
    77.8%
    Agentic Performance %
    52.1%
    Arena Preference Elo
    839
    Task Success Rate %
    80.7%
    Time to First Audio Seconds
    1.25
Compare this task on the human ranking →

Music generation

For generating complete instrumental or vocal music.

Approximately tied at the top. Instrumental and vocal preference remain separate lists. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
audio
Metric / benchmark
Music Arena — Instrumental
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Suno V5.5
Consensus #2
Mureka V9
Consensus #3
Mureka V8
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Artificial Analysis":1}

Instrumental

  1. #1 Suno V5.5

    Base model
    Suno V5.5
    Provider
    Suno
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisMusic Arena — Instrumental
    Source model name
    Suno V5.5
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 1
    Raw score
    1186
    Unit
    Elo
    Confidence interval
    1177 / 1195
    Sample count
    6469
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:29.413Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Mureka V9

    Base model
    Mureka V9
    Provider
    Mureka
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisMusic Arena — Instrumental
    Source model name
    Mureka V9
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 3
    Raw score
    1177
    Unit
    Elo
    Confidence interval
    1163 / 1191
    Sample count
    3046
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:29.413Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Mureka V8

    Base model
    Mureka V8
    Provider
    Mureka
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisMusic Arena — Instrumental
    Source model name
    Mureka V8
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    3 / 4
    Raw score
    1166
    Unit
    Elo
    Confidence interval
    1158 / 1174
    Sample count
    7497
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:29.413Z
    Evidence freshness
    unknown
    Used in aggregation
    true

Vocals

  1. #1 Suno V5.5

    Base model
    Suno V5.5
    Provider
    Suno
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisMusic Arena — Vocals
    Source model name
    Suno V5.5
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 1
    Raw score
    1170
    Unit
    Elo
    Confidence interval
    1163 / 1177
    Sample count
    10373
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:29.205Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Mureka V9

    Base model
    Mureka V9
    Provider
    Mureka
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisMusic Arena — Vocals
    Source model name
    Mureka V9
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 2
    Raw score
    1159
    Unit
    Elo
    Confidence interval
    1145 / 1173
    Sample count
    3010
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:29.205Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Mureka V8

    Base model
    Mureka V8
    Provider
    Mureka
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisMusic Arena — Vocals
    Source model name
    Mureka V8
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    3 / 3
    Raw score
    1140
    Unit
    Elo
    Confidence interval
    1134 / 1146
    Sample count
    11414
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:29.205Z
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Data analysis

For objective data-analysis problems; dashboard design and standalone SQL quality are separate tasks.

Objective data analysis; this is not a web dashboard design benchmark and does not establish a separate SQL winner. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
code
Metric / benchmark
LiveBench Data Analysis
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
GPT-6 Astra
Consensus #2
GPT-5.5
Consensus #3
Claude Fable 5
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
clear_winner
CI overlap
false
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"LiveBench":1}

Top 3

  1. #1 GPT-6 Astra

    Base model
    GPT-6 Astra
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    LiveBenchLiveBench Data Analysis
    Source model name
    GPT-6 Astra Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    83
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 GPT-5.5

    Base model
    GPT-5.5
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    LiveBenchLiveBench Data Analysis
    Source model name
    GPT-5.5 Thinking xHigh Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    81.6
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    LiveBenchLiveBench Data Analysis
    Source model name
    Claude Fable 5 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    80.5
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Structured output and tool calling

For reliable function calls, JSON arguments and structured tool use.

BFCL function-calling accuracy, with explicit benchmark configurations. Single source family; independent cross-source confirmation is unavailable in this edition. Warning: the published benchmark snapshot is more than 30 days old.

Category
code
Metric / benchmark
BFCL overall function calling
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude-Opus-4.5-20251101
Consensus #2
Claude-Sonnet-4.5-20250929
Consensus #3
Gemini-3-Pro-Preview
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
low
Disagreement status
clear_winner
CI overlap
false
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Berkeley BFCL":1}

Top 3

  1. #1 Claude-Opus-4.5-20251101

    Base model
    Claude-Opus-4.5-20251101
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    function calling
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    low
    Source-specific results
    Berkeley BFCLBFCL overall function calling
    Source model name
    Claude-Opus-4-5-20251101 (FC)
    Benchmark version
    V4 / f7cf735
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    77.47
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    function calling
    Reasoning effort
    Not reported
    Source date
    2026-04-12
    Captured at
    2026-09-07T10:10:33.001Z
    Evidence freshness
    stale
    Used in aggregation
    true
    Berkeley BFCLBFCL overall function calling
    Source model name
    Claude-Opus-4-5-20251101 (Prompt)
    Benchmark version
    V4 / f7cf735
    Source rank
    57
    Rank spread
    Not reported / Not reported
    Raw score
    33.47
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    function calling
    Reasoning effort
    Not reported
    Source date
    2026-04-12
    Captured at
    2026-09-07T10:10:33.001Z
    Evidence freshness
    stale
    Used in aggregation
    false
  2. #2 Claude-Sonnet-4.5-20250929

    Base model
    Claude-Sonnet-4.5-20250929
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    function calling
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    low
    Source-specific results
    Berkeley BFCLBFCL overall function calling
    Source model name
    Claude-Sonnet-4-5-20250929 (FC)
    Benchmark version
    V4 / f7cf735
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    73.24
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    function calling
    Reasoning effort
    Not reported
    Source date
    2026-04-12
    Captured at
    2026-09-07T10:10:33.001Z
    Evidence freshness
    stale
    Used in aggregation
    true
    Berkeley BFCLBFCL overall function calling
    Source model name
    Claude-Sonnet-4-5-20250929 (Prompt)
    Benchmark version
    V4 / f7cf735
    Source rank
    89
    Rank spread
    Not reported / Not reported
    Raw score
    24.9
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    function calling
    Reasoning effort
    Not reported
    Source date
    2026-04-12
    Captured at
    2026-09-07T10:10:33.001Z
    Evidence freshness
    stale
    Used in aggregation
    false
  3. #3 Gemini-3-Pro-Preview

    Base model
    Gemini-3-Pro-Preview
    Provider
    Google
    Official model ID
    Not reported
    Variant
    function calling
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    low
    Source-specific results
    Berkeley BFCLBFCL overall function calling
    Source model name
    Gemini-3-Pro-Preview (Prompt)
    Benchmark version
    V4 / f7cf735
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    72.51
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    function calling
    Reasoning effort
    Not reported
    Source date
    2026-04-12
    Captured at
    2026-09-07T10:10:33.001Z
    Evidence freshness
    stale
    Used in aggregation
    true
    Berkeley BFCLBFCL overall function calling
    Source model name
    Gemini-3-Pro-Preview (FC)
    Benchmark version
    V4 / f7cf735
    Source rank
    7
    Rank spread
    Not reported / Not reported
    Raw score
    68.14
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    function calling
    Reasoning effort
    Not reported
    Source date
    2026-04-12
    Captured at
    2026-09-07T10:10:33.001Z
    Evidence freshness
    stale
    Used in aggregation
    false
Compare this task on the human ranking →

Controlled-voice synthesis

For comparing speech synthesis with a controlled common voice; speaker similarity is not evaluated.

Approximately tied at the top. Controlled Voice Arena tests common-voice synthesis preference; it does not separately establish speaker identity similarity. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
audio
Metric / benchmark
Controlled Voice Arena
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Realtime TTS-2
Consensus #2
Sonic 3.6
Consensus #3
Sonic 3.5
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
low
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Artificial Analysis":1}

Top 3

  1. #1 Realtime TTS-2

    Base model
    Realtime TTS-2
    Provider
    Inworld
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    low
    Source-specific results
    Artificial AnalysisControlled Voice Arena
    Source model name
    Realtime TTS-2
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 2
    Raw score
    1123
    Unit
    Elo
    Confidence interval
    1107 / 1139
    Sample count
    1292
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:32.121Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Sonic 3.6

    Base model
    Sonic 3.6
    Provider
    Cartesia
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    low
    Source-specific results
    Artificial AnalysisControlled Voice Arena
    Source model name
    Sonic 3.6
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 2
    Raw score
    1119
    Unit
    Elo
    Confidence interval
    1104 / 1134
    Sample count
    1993
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:32.121Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Sonic 3.5

    Base model
    Sonic 3.5
    Provider
    Cartesia
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    low
    Source-specific results
    Artificial AnalysisControlled Voice Arena
    Source model name
    Sonic 3.5
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    3 / 3
    Raw score
    1096
    Unit
    Elo
    Confidence interval
    1084 / 1108
    Sample count
    3791
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:32.121Z
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Code reasoning

Objective coding problems without a terminal agent harness.

Objective coding problems without a terminal agent harness. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
code
Metric / benchmark
LiveBench Coding
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5.1
Consensus #2
Claude Fable 5
Consensus #3
GPT-5.6 Sol
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
clear_winner
CI overlap
false
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"LiveBench":1}

Top 3

  1. #1 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    LiveBenchLiveBench Coding
    Source model name
    Claude Fable 5.1 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    86.4
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    LiveBenchLiveBench Coding
    Source model name
    Claude Fable 5 Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    86
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 GPT-5.6 Sol

    Base model
    GPT-5.6 Sol
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    LiveBenchLiveBench Coding
    Source model name
    GPT-5.6 Sol Max Effort
    Benchmark version
    2026-06-25 (latest release)
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    83.9
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:19.053Z
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Reference-based frontend design

Recreate a frontend from a design reference; separate from general WebDev.

Approximately tied at the top. Recreate a frontend from a design reference; separate from general WebDev. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
code
Metric / benchmark
Code Arena | WebDev🎨Reference-Based Design
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5.1
Consensus #2
GPT-6 Astra
Consensus #3
Kimi K3 Max
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Arena":1}

Top 3

  1. #1 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaCode Arena | WebDev🎨Reference-Based Design
    Source model name
    claude-fable-5.1-max
    Benchmark version
    live leaderboard
    Source rank
    1
    Rank spread
    1 / 2
    Raw score
    1815
    Unit
    Elo
    Confidence interval
    1773 / 1857
    Sample count
    366
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
  2. #2 GPT-6 Astra

    Base model
    GPT-6 Astra
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.63092975
    Mean rank
    2
    Median rank
    2
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaCode Arena | WebDev🎨Reference-Based Design
    Source model name
    gpt-6-astra-max
    Benchmark version
    live leaderboard
    Source rank
    2
    Rank spread
    1 / 3
    Raw score
    1785
    Unit
    Elo
    Confidence interval
    1731 / 1839
    Sample count
    220
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
  3. #3 Kimi K3 Max

    Base model
    Kimi K3 Max
    Provider
    Moonshot
    Official model ID
    Not reported
    Variant
    Max
    Reasoning effort
    max
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.5
    Mean rank
    3
    Median rank
    3
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    1
    Confidence
    medium
    Source-specific results
    ArenaCode Arena | WebDev🎨Reference-Based Design
    Source model name
    kimi-k3-max
    Benchmark version
    live leaderboard
    Source rank
    3
    Rank spread
    2 / 8
    Raw score
    1723
    Unit
    Elo
    Confidence interval
    1695 / 1751
    Sample count
    777
    Preliminary
    false
    Variant
    Max
    Reasoning effort
    max
    Source date
    2026-09-05
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    fresh
    Used in aggregation
    true
Compare this task on the human ranking →

Repository refactoring

Restructure repositories while preserving behavior, using the stated agent harness.

Approximately tied at the top. Restructure repositories while preserving behavior, using the stated agent harness. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
code
Metric / benchmark
SWE Atlas — Refactoring
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Claude Fable 5.1
Consensus #2
Claude Fable 5
Consensus #3
Claude Opus 4.7
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
statistical_tie
CI overlap
true
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Scale Labs":1}

Top 3

  1. #1 Claude Fable 5.1

    Base model
    Claude Fable 5.1
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Scale LabsSWE Atlas — Refactoring
    Source model name
    Fable-5.1 (Claude Code) xHigh
    Benchmark version
    current leaderboard
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    56.67
    Unit
    percent
    Confidence interval
    50.150000000000006 / 63.19
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  2. #2 Claude Fable 5

    Base model
    Claude Fable 5
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Scale LabsSWE Atlas — Refactoring
    Source model name
    Fable-5 (Claude Code) xHigh
    Benchmark version
    current leaderboard
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    54.76
    Unit
    percent
    Confidence interval
    48 / 61.519999999999996
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Xhigh
    Reasoning effort
    xhigh
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
  3. #3 Claude Opus 4.7

    Base model
    Claude Opus 4.7
    Provider
    Anthropic
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Scale LabsSWE Atlas — Refactoring
    Source model name
    Opus-4.7 (Claude Code)
    Benchmark version
    current leaderboard
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    48.57
    Unit
    percent
    Confidence interval
    41.84 / 55.3
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:16:44.291125+00:00
    Evidence freshness
    unknown
    Used in aggregation
    true
Compare this task on the human ranking →

Realtime voice conversation

Equal rank weight for speech reasoning and conversational dynamics. The composite index and agent completion are retained as component context.

Sources disagree on ranking order. Equal rank weight for speech reasoning and conversational dynamics. The composite index and agent completion are retained as component context. Single source family; independent cross-source confirmation is unavailable in this edition.

Category
audio
Metric / benchmark
Speech Reasoning; Conversational Dynamics
Consensus metric direction
higher_is_better
Raw metric direction
higher_is_better
Consensus #1
Qwen Audio 3.0 Realtime Plus
Consensus #2
Qwen Audio 3.0 Realtime Flash
Consensus #3
GPT-Realtime-2
Consensus score
1
Mean rank
1
Median rank
1
Independent source count
1
Source coverage
1
Confidence
medium
Disagreement status
source_split
CI overlap
false
Last verified
2026-09-07T10:16:44.291125+00:00
Source family weights
{"Artificial Analysis":1}

Top 3

  1. #1 Qwen Audio 3.0 Realtime Plus

    Base model
    Qwen Audio 3.0 Realtime Plus
    Provider
    Alibaba Cloud
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    1
    Evaluation status
    evaluated
    Consensus score
    1
    Mean rank
    1
    Median rank
    1
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisSpeech Reasoning
    Source model name
    Qwen Audio 3.0 Realtime Plus
    Benchmark version
    current components
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    99
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Speech to Speech Index %
    66.8%
    Speech Reasoning %
    99%
    Conversational Dynamics %
    98.4%
    Agentic Performance %
    54.6%
    Arena Preference Elo
    699
    Task Success Rate %
    77.8%
    Time to First Audio Seconds
    1.54
    Artificial AnalysisConversational Dynamics
    Source model name
    Qwen Audio 3.0 Realtime Plus
    Benchmark version
    current components
    Source rank
    1
    Rank spread
    Not reported / Not reported
    Raw score
    98.4
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Speech to Speech Index %
    66.8%
    Speech Reasoning %
    99%
    Conversational Dynamics %
    98.4%
    Agentic Performance %
    54.6%
    Arena Preference Elo
    699
    Task Success Rate %
    77.8%
    Time to First Audio Seconds
    1.54
  2. #2 Qwen Audio 3.0 Realtime Flash

    Base model
    Qwen Audio 3.0 Realtime Flash
    Provider
    Alibaba Cloud
    Official model ID
    Not reported
    Variant
    Not reported
    Reasoning effort
    Not reported
    Consensus rank
    2
    Evaluation status
    evaluated
    Consensus score
    0.47319732
    Mean rank
    5
    Median rank
    5
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisSpeech Reasoning
    Source model name
    Qwen Audio 3.0 Realtime Flash
    Benchmark version
    current components
    Source rank
    8
    Rank spread
    Not reported / Not reported
    Raw score
    96
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Speech to Speech Index %
    64.2%
    Speech Reasoning %
    96%
    Conversational Dynamics %
    96.9%
    Agentic Performance %
    35.9%
    Arena Preference Elo
    752
    Task Success Rate %
    81.7%
    Time to First Audio Seconds
    1.55
    Artificial AnalysisConversational Dynamics
    Source model name
    Qwen Audio 3.0 Realtime Flash
    Benchmark version
    current components
    Source rank
    2
    Rank spread
    Not reported / Not reported
    Raw score
    96.9
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Not reported
    Reasoning effort
    Not reported
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Speech to Speech Index %
    64.2%
    Speech Reasoning %
    96%
    Conversational Dynamics %
    96.9%
    Agentic Performance %
    35.9%
    Arena Preference Elo
    752
    Task Success Rate %
    81.7%
    Time to First Audio Seconds
    1.55
  3. #3 GPT-Realtime-2

    Base model
    GPT-Realtime-2
    Provider
    OpenAI
    Official model ID
    Not reported
    Variant
    High
    Reasoning effort
    high
    Consensus rank
    3
    Evaluation status
    evaluated
    Consensus score
    0.46533828
    Mean rank
    3.5
    Median rank
    3.5
    Independent source count
    1
    Source coverage
    1
    Fresh source count
    0
    Confidence
    medium
    Source-specific results
    Artificial AnalysisSpeech Reasoning
    Source model name
    GPT-Realtime-2 (High)
    Benchmark version
    current components
    Source rank
    4
    Rank spread
    Not reported / Not reported
    Raw score
    97
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Speech to Speech Index %
    73.6%
    Speech Reasoning %
    97%
    Conversational Dynamics %
    95.3%
    Agentic Performance %
    39.8%
    Arena Preference Elo
    914
    Task Success Rate %
    89.8%
    Time to First Audio Seconds
    1.14
    Artificial AnalysisSpeech Reasoning
    Source model name
    GPT-Realtime-2 (Medium)
    Benchmark version
    current components
    Source rank
    10
    Rank spread
    Not reported / Not reported
    Raw score
    93
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Medium
    Reasoning effort
    medium
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Speech to Speech Index %
    -
    Speech Reasoning %
    93%
    Conversational Dynamics %
    95.2%
    Agentic Performance %
    37.4%
    Arena Preference Elo
    -
    Task Success Rate %
    -
    Time to First Audio Seconds
    1.22
    Artificial AnalysisSpeech Reasoning
    Source model name
    GPT-Realtime-2 (Minimal)
    Benchmark version
    current components
    Source rank
    19
    Rank spread
    Not reported / Not reported
    Raw score
    72
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Minimal
    Reasoning effort
    minimal
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Speech to Speech Index %
    62.7%
    Speech Reasoning %
    72%
    Conversational Dynamics %
    96.1%
    Agentic Performance %
    30.8%
    Arena Preference Elo
    881
    Task Success Rate %
    84.7%
    Time to First Audio Seconds
    1.12
    Artificial AnalysisConversational Dynamics
    Source model name
    GPT-Realtime-2 (High)
    Benchmark version
    current components
    Source rank
    7
    Rank spread
    Not reported / Not reported
    Raw score
    95.3
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    High
    Reasoning effort
    high
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Speech to Speech Index %
    73.6%
    Speech Reasoning %
    97%
    Conversational Dynamics %
    95.3%
    Agentic Performance %
    39.8%
    Arena Preference Elo
    914
    Task Success Rate %
    89.8%
    Time to First Audio Seconds
    1.14
    Artificial AnalysisConversational Dynamics
    Source model name
    GPT-Realtime-2 (Medium)
    Benchmark version
    current components
    Source rank
    8
    Rank spread
    Not reported / Not reported
    Raw score
    95.2
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Medium
    Reasoning effort
    medium
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    false
    Speech to Speech Index %
    -
    Speech Reasoning %
    93%
    Conversational Dynamics %
    95.2%
    Agentic Performance %
    37.4%
    Arena Preference Elo
    -
    Task Success Rate %
    -
    Time to First Audio Seconds
    1.22
    Artificial AnalysisConversational Dynamics
    Source model name
    GPT-Realtime-2 (Minimal)
    Benchmark version
    current components
    Source rank
    3
    Rank spread
    Not reported / Not reported
    Raw score
    96.1
    Unit
    percent
    Confidence interval
    Not reported / Not reported
    Sample count
    Not reported
    Preliminary
    false
    Variant
    Minimal
    Reasoning effort
    minimal
    Source date
    Not reported
    Captured at
    2026-09-07T10:10:23.006Z
    Evidence freshness
    unknown
    Used in aggregation
    true
    Speech to Speech Index %
    62.7%
    Speech Reasoning %
    72%
    Conversational Dynamics %
    96.1%
    Agentic Performance %
    30.8%
    Arena Preference Elo
    881
    Task Success Rate %
    84.7%
    Time to First Audio Seconds
    1.12
Compare this task on the human ranking →

Tasks awaiting comparable evidence

Translation: unavailable. Previous ranking used an agency article and unnamed model families; comparable version-specific primary evidence has not been established.