Human-oriented ranking Raw ranking data: JSON Task-specific AI model rankings — structured for agents and search systems. Each task has its own top models, evidence, source dates, confidence and disagreement notes. Newer does not automatically mean better.
Last verified: 2026-09-07. All rankings are derived from task-relevant public or independent evidence.
29 tasks · 39 ranked base models · Methodology 2.0-task-consensus
General chat Russian language Expert analysis Mathematics Instruction following Copywriting Search and research PDF and documents Terminal software engineering Web development AI agents Image understanding Text from image (OCR) Image generation Image editing Text → video Image → video Video editing Speech to text Text to speech Voice agents Music generation Data analysis Structured output and tool calling Controlled-voice synthesis Code reasoning Reference-based frontend design Repository refactoring Realtime voice conversation Consensus methodology Consensus rankings use only task-relevant sources. Source rankings are normalized within each benchmark before aggregation. Raw scores from different benchmark families are never averaged. Vendor-published benchmarks do not determine rank unless independently reproduced. Models that have not yet been evaluated remain pending rather than being assigned an artificial low score.
method: weighted_reciprocal_rank
formula: weighted_mean(1 / log2(rank + 1))
raw scores averaged: false
source family policy: Within each source family, average benchmark reciprocal ranks; then weight source families. Republished benchmark results are not independent.
configuration policy: One base model per top three. Best reported rank configuration per benchmark contributes once; every raw configuration remains in source_results. Different versions remain separate base models. H3 Max is explicitly a fal post-trained configuration.
admission policy: Final results take priority over preliminary ones. When at least three models have two independent source families, only those enter the consensus top three. Other evaluated models remain candidates; missing evidence is never a bad score.
tie break: Exact consensus ties: coverage, then mean displayed source position, then stable model_id; this order is not statistical superiority.
freshness policy: published_at is the dated leaderboard snapshot, never a model release date or benchmark release version. Unknown publication dates stay null. Evidence older than 30 days is stale; 14 days is a warning.
limitations: Source coverage and confidence describe this captured evidence set. No vendor-reported result determines rank. Unknown API IDs, prices and context sizes are not inferred.
General chat For broad conversational quality across everyday prompts.
Sources disagree on #1. Human preference is weighted 70%; objective broad ability provides two secondary checks.
Category text
Metric / benchmark Text Arena🏆Overall; LiveBench Overall; Intelligence Index
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5
Consensus #2 Claude Fable 5.1
Consensus #3 GPT-6 Astra
Consensus score 0.84463946
Mean rank 3.3333333333333335
Median rank 2
Independent source count 3
Source coverage 1
Confidence medium
Disagreement status source_split
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":0.7,"Artificial Analysis":0.15,"LiveBench":0.15} Top 3 #1 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 0.84463946
Mean rank 3.3333333333333335
Median rank 2
Independent source count 3
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 6
Raw score 1507
Unit Elo
Confidence interval 1502 / 1512
Sample count 27189
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Claude Fable 5 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 2
Rank spread Not reported / Not reported
Raw score 83
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true
Source model name Claude Fable 5 (with fallback)
Benchmark version 4.2
Source rank 7
Rank spread Not reported / Not reported
Raw score 53
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation true #2 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 2
Evaluation status evaluated
Consensus score 0.65
Mean rank 1.6666666666666667
Median rank 1
Independent source count 3
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5.1-max
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 14
Raw score 1504
Unit Elo
Confidence interval 1493 / 1515
Sample count 2906
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Claude Fable 5.1 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 1
Rank spread Not reported / Not reported
Raw score 83.4
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true
Source model name Claude Fable 5.1 (max with fallback)
Benchmark version 4.2
Source rank 1
Rank spread Not reported / Not reported
Raw score 57
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation true
Source model name Claude Fable 5.1 (xhigh with fallback)
Benchmark version 4.2
Source rank 2
Rank spread Not reported / Not reported
Raw score 56
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Xhigh
Reasoning effort xhigh
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation false
Source model name Claude Fable 5.1 (high with fallback)
Benchmark version 4.2
Source rank 4
Rank spread Not reported / Not reported
Raw score 54
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation false
Source model name Claude Fable 5.1 (medium with fallback)
Benchmark version 4.2
Source rank 7
Rank spread Not reported / Not reported
Raw score 53
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Medium
Reasoning effort medium
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation false
Source model name Claude Fable 5.1 (low with fallback)
Benchmark version 4.2
Source rank 15
Rank spread Not reported / Not reported
Raw score 51
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Low
Reasoning effort low
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation false #3 GPT-6 Astra
Base model GPT-6 Astra
Provider OpenAI
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 3
Evaluation status partially_evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 2
Source coverage 0.6666666666666666
Fresh source count 0
Confidence medium Source-specific results
Source model name GPT-6 Astra Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 3
Rank spread Not reported / Not reported
Raw score 82.2
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true
Source model name GPT-6 Astra (max)
Benchmark version 4.2
Source rank 3
Rank spread Not reported / Not reported
Raw score 55
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation true
Source model name GPT-6 Astra (xhigh)
Benchmark version 4.2
Source rank 4
Rank spread Not reported / Not reported
Raw score 54
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Xhigh
Reasoning effort xhigh
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation false
Source model name GPT-6 Astra (high)
Benchmark version 4.2
Source rank 7
Rank spread Not reported / Not reported
Raw score 53
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation false
Source model name GPT-6 Astra (medium)
Benchmark version 4.2
Source rank 12
Rank spread Not reported / Not reported
Raw score 52
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Medium
Reasoning effort medium
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation false
Source model name GPT-6 Astra (low)
Benchmark version 4.2
Source rank 22
Rank spread Not reported / Not reported
Raw score 49
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Low
Reasoning effort low
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation false
Source model name GPT-6 Astra (Non-reasoning)
Benchmark version 4.2
Source rank 25
Rank spread Not reported / Not reported
Raw score 48
Unit index
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:22.585Z
Evidence freshness unknown
Used in aggregation false Compare this task on the human ranking → Russian language For Russian-language conversation and writing.
Approximately tied at the top. Russian-language preference only; overlapping rank ranges limit certainty. Single source family; independent cross-source confirmation is unavailable in this edition.
Category text
Metric / benchmark Text Arena🇷🇺Russian
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5
Consensus #2 Muse Spark 1.2
Consensus #3 Claude Fable 5.1
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1} Top 3 #1 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 11
Raw score 1521
Unit Elo
Confidence interval 1509 / 1533
Sample count 2706
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true #2 Muse Spark 1.2
Base model Muse Spark 1.2
Provider Meta
Official model ID Not reported
Variant Xhigh
Reasoning effort xhigh
Consensus rank 2
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name muse-spark-1.2 (xHigh)
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 39
Raw score 1514
Unit Elo
Confidence interval 1482 / 1546
Sample count 330
Preliminary false
Variant Xhigh
Reasoning effort xhigh
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true #3 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 3
Evaluation status evaluated
Consensus score 0.43067656
Mean rank 4
Median rank 4
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5.1-max
Benchmark version live leaderboard
Source rank 4
Rank spread 1 / 46
Raw score 1510
Unit Elo
Confidence interval 1477 / 1543
Sample count 348
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true Compare this task on the human ranking → Expert analysis For difficult analytical and professional knowledge tasks.
Sources disagree on #1. General expert reasoning; does not establish a winner for individual professions.
Category text
Metric / benchmark Text Arena🤓Expert; LiveBench Reasoning; Humanity’s Last Exam
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5
Consensus #2 Claude Fable 5.1
Consensus #3 Claude Opus 4.6
Consensus score 0.73788625
Mean rank 5
Median rank 5
Independent source count 3
Source coverage 0.6666666666666666
Confidence medium
Disagreement status source_split
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":0.5,"LiveBench":0.3,"Scale Labs":0.2} Top 3 #1 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status partially_evaluated
Consensus score 0.73788625
Mean rank 5
Median rank 5
Independent source count 2
Source coverage 0.6666666666666666
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 13
Raw score 1548
Unit Elo
Confidence interval 1537 / 1559
Sample count 2939
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Claude Fable 5 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 9
Rank spread Not reported / Not reported
Raw score 89.7
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true #2 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 2
Evaluation status evaluated
Consensus score 0.55594559
Mean rank 3.3333333333333335
Median rank 2
Independent source count 3
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5.1-max
Benchmark version live leaderboard
Source rank 7
Rank spread 1 / 65
Raw score 1532
Unit Elo
Confidence interval 1490 / 1574
Sample count 211
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Claude Fable 5.1 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 2
Rank spread Not reported / Not reported
Raw score 91.7
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true
Source model name Fable 5.1 (xhigh)
Benchmark version current leaderboard
Source rank 1
Rank spread Not reported / Not reported
Raw score 46.5
Unit percent
Confidence interval 44.5 / 48.5
Sample count Not reported
Preliminary false
Variant Xhigh
Reasoning effort xhigh
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 Claude Opus 4.6
Base model Claude Opus 4.6
Provider Anthropic
Official model ID Not reported
Variant High
Reasoning effort high
Consensus rank 3
Evaluation status partially_evaluated
Consensus score 0.52056426
Mean rank 9
Median rank 9
Independent source count 2
Source coverage 0.6666666666666666
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-opus-4-6-high
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 13
Raw score 1547
Unit Elo
Confidence interval 1538 / 1556
Sample count 6409
Preliminary false
Variant High
Reasoning effort high
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name claude-opus-4-6
Benchmark version live leaderboard
Source rank 5
Rank spread 1 / 18
Raw score 1534
Unit Elo
Confidence interval 1526 / 1542
Sample count 7644
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation false
Source model name claude-opus-4-6 (Non-Thinking)
Benchmark version current leaderboard
Source rank 16
Rank spread Not reported / Not reported
Raw score 19
Unit percent
Confidence interval 17.46 / 20.54
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Mathematics For hard calculations, proofs and multi-step maths.
Sources disagree on #1. Mathematics receives 60%; HLE is a broader reasoning cross-check, not a math-only score.
Category text
Metric / benchmark Text Arena🧮Math; LiveBench Mathematics; Humanity’s Last Exam
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5.1
Consensus #2 Claude Fable 5
Consensus #3 Claude Opus 5
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 3
Source coverage 0.6666666666666666
Confidence medium
Disagreement status source_split
CI overlap false
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":0.25,"LiveBench":0.6,"Scale Labs":0.15} Top 3 #1 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 1
Evaluation status partially_evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 2
Source coverage 0.6666666666666666
Fresh source count 0
Confidence medium Source-specific results
Source model name Claude Fable 5.1 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 1
Rank spread Not reported / Not reported
Raw score 97
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true
Source model name Fable 5.1 (xhigh)
Benchmark version current leaderboard
Source rank 1
Rank spread Not reported / Not reported
Raw score 46.5
Unit percent
Confidence interval 44.5 / 48.5
Sample count Not reported
Preliminary false
Variant Xhigh
Reasoning effort xhigh
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status partially_evaluated
Consensus score 0.59812463
Mean rank 2.5
Median rank 2.5
Independent source count 2
Source coverage 0.6666666666666666
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 15
Raw score 1529
Unit Elo
Confidence interval 1513 / 1545
Sample count 1331
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Claude Fable 5 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 4
Rank spread Not reported / Not reported
Raw score 96
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true #3 Claude Opus 5
Base model Claude Opus 5
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 3
Evaluation status partially_evaluated
Consensus score 0.42086169
Mean rank 4.5
Median rank 4.5
Independent source count 2
Source coverage 0.6666666666666666
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-opus-5-max
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 18
Raw score 1529
Unit Elo
Confidence interval 1507 / 1551
Sample count 740
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name claude-opus-5-high
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 16
Raw score 1526
Unit Elo
Confidence interval 1510 / 1542
Sample count 1523
Preliminary false
Variant High
Reasoning effort high
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation false
Source model name Claude 5 Opus Thinking Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 7
Rank spread Not reported / Not reported
Raw score 95.7
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Instruction following For prompts with detailed constraints and required formats.
Sources disagree on #1. Preference, objective instruction following and IFBench measure different constraints; see each result.
Category text
Metric / benchmark Text Arena📝Instruction Following; LiveBench Instruction Following; IFBench
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Grok 4.3
Consensus #2 Claude Fable 5
Consensus #3 Gemini 3.1 Pro Preview
Consensus score 0.44489779
Mean rank 49.333333333333336
Median rank 40
Independent source count 3
Source coverage 1
Confidence medium
Disagreement status source_split
CI overlap false
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1,"Artificial Analysis":1,"LiveBench":1} Top 3 #1 Grok 4.3
Base model Grok 4.3
Provider SpaceXAI
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 0.44489779
Mean rank 49.333333333333336
Median rank 40
Independent source count 3
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name grok-4.3
Benchmark version live leaderboard
Source rank 107
Rank spread 86 / 130
Raw score 1416
Unit Elo
Confidence interval 1411 / 1421
Sample count 22801
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Grok 4.3
Benchmark version 2026-06-25 (latest release)
Source rank 40
Rank spread Not reported / Not reported
Raw score 62.8
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true
Source model name Grok 4.3 (medium)
Benchmark version 58 constraints; visible top-10 chart
Source rank 1
Rank spread Not reported / Not reported
Raw score 83.3
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Medium
Reasoning effort medium
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true
Source model name Grok 4.3 (high)
Benchmark version 58 constraints; visible top-10 chart
Source rank 5
Rank spread Not reported / Not reported
Raw score 81.3
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation false #2 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.42540059
Mean rank 6
Median rank 6
Independent source count 3
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 7
Raw score 1512
Unit Elo
Confidence interval 1505 / 1519
Sample count 9607
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Claude Fable 5 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 6
Rank spread Not reported / Not reported
Raw score 75.8
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true
Source model name Claude Fable 5 (with fallback)
Benchmark version 58 constraints; visible top-10 chart
Source rank 10
Rank spread Not reported / Not reported
Raw score 63.5
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 Gemini 3.1 Pro Preview
Base model Gemini 3.1 Pro Preview
Provider Google
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status partially_evaluated
Consensus score 0.375
Mean rank 9
Median rank 9
Independent source count 2
Source coverage 0.6666666666666666
Fresh source count 1
Confidence medium Source-specific results
Source model name gemini-3.1-pro-preview
Benchmark version live leaderboard
Source rank 15
Rank spread 9 / 31
Raw score 1481
Unit Elo
Confidence interval 1476 / 1486
Sample count 34766
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Gemini 3.1 Pro Preview High
Benchmark version 2026-06-25 (latest release)
Source rank 3
Rank spread Not reported / Not reported
Raw score 79.1
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Copywriting For persuasive, creative and polished marketing text.
Sources disagree on #1. Creative writing and writing/literature/language without style control. These are proxies for copywriting, not conversion tests. Single source family; independent cross-source confirmation is unavailable in this edition.
Category text
Metric / benchmark Text Arena✍️Creative Writing; Text Arena✍️Writing, Literature, & Language
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5
Consensus #2 Claude Fable 5.1
Consensus #3 Claude Opus 4.6
Consensus score 0.81546488
Mean rank 1.5
Median rank 1.5
Independent source count 1
Source coverage 1
Confidence low
Disagreement status source_split
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1} Top 3 #1 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 0.81546488
Mean rank 1.5
Median rank 1.5
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence low Source-specific results
Source model name claude-fable-5
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 6
Raw score 1504
Unit Elo
Confidence interval 1495 / 1513
Sample count 5480
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name claude-fable-5
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 8
Raw score 1505
Unit Elo
Confidence interval 1497 / 1513
Sample count 7328
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true #2 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 2
Evaluation status evaluated
Consensus score 0.67810359
Mean rank 3.5
Median rank 3.5
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence low Source-specific results
Source model name claude-fable-5.1-max
Benchmark version live leaderboard
Source rank 6
Rank spread 1 / 30
Raw score 1487
Unit Elo
Confidence interval 1462 / 1512
Sample count 608
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name claude-fable-5.1-max
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 7
Raw score 1526
Unit Elo
Confidence interval 1504 / 1548
Sample count 782
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true #3 Claude Opus 4.6
Base model Claude Opus 4.6
Provider Anthropic
Official model ID Not reported
Variant High
Reasoning effort high
Consensus rank 3
Evaluation status evaluated
Consensus score 0.56546488
Mean rank 2.5
Median rank 2.5
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence low Source-specific results
Source model name claude-opus-4-6-high
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 6
Raw score 1500
Unit Elo
Confidence interval 1493 / 1507
Sample count 12678
Preliminary false
Variant High
Reasoning effort high
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name claude-opus-4-6
Benchmark version live leaderboard
Source rank 9
Rank spread 3 / 22
Raw score 1479
Unit Elo
Confidence interval 1472 / 1486
Sample count 12724
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation false
Source model name claude-opus-4-6-high
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 9
Raw score 1502
Unit Elo
Confidence interval 1496 / 1508
Sample count 18186
Preliminary false
Variant High
Reasoning effort high
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name claude-opus-4-6
Benchmark version live leaderboard
Source rank 7
Rank spread 2 / 10
Raw score 1494
Unit Elo
Confidence interval 1488 / 1500
Sample count 18365
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation false Compare this task on the human ranking → Search and research For web search, evidence gathering and sourced answers.
Approximately tied at the top. Search-enabled configurations only. General intelligence is not search evidence. Single source family; independent cross-source confirmation is unavailable in this edition.
Category text
Metric / benchmark Search Arena
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 GPT-5.6 Sol
Consensus #2 Claude Opus 4.6
Consensus #3 GPT-5.5
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1} Top 3 #1 GPT-5.6 Sol
Base model GPT-5.6 Sol
Provider OpenAI
Official model ID Not reported
Variant Xhigh
Reasoning effort xhigh
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name gpt-5.6-sol-xhigh
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 2
Raw score 1257
Unit Elo
Confidence interval 1250 / 1264
Sample count 29663
Preliminary false
Variant Xhigh
Reasoning effort xhigh
Source date 2026-08-24
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness warning
Used in aggregation true #2 Claude Opus 4.6
Base model Claude Opus 4.6
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name claude-opus-4-6-search
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 2
Raw score 1253
Unit Elo
Confidence interval 1248 / 1258
Sample count 134699
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-08-24
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness warning
Used in aggregation true #3 GPT-5.5
Base model GPT-5.5
Provider OpenAI
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name gpt-5.5-search
Benchmark version live leaderboard
Source rank 3
Rank spread 3 / 5
Raw score 1242
Unit Elo
Confidence interval 1237 / 1247
Sample count 89873
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-08-24
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness warning
Used in aggregation true Pending evaluation GPT-6 Astra: pending_evaluation. New release — awaiting comparable task-level evidence. Claude Fable 5.1: pending_evaluation. New release — awaiting comparable task-level evidence. Compare this task on the human ranking → PDF and documents For understanding long documents, PDFs and mixed document content.
Approximately tied at the top. Document Arena preference; newly released models without a document evaluation remain pending. Single source family; independent cross-source confirmation is unavailable in this edition.
Category text
Metric / benchmark Document Arena
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Opus 5
Consensus #2 Claude Opus 4.6
Consensus #3 Claude Fable 5
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1} Top 3 #1 Claude Opus 5
Base model Claude Opus 5
Provider Anthropic
Official model ID Not reported
Variant High
Reasoning effort high
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name claude-opus-5-high
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 4
Raw score 1520
Unit Elo
Confidence interval 1505 / 1535
Sample count 1663
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 Claude Opus 4.6
Base model Claude Opus 4.6
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name claude-opus-4-6
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 6
Raw score 1510
Unit Elo
Confidence interval 1504 / 1516
Sample count 37271
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true
Source model name claude-opus-4-6-high
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 6
Raw score 1506
Unit Elo
Confidence interval 1499 / 1513
Sample count 24973
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation false #3 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.43067656
Mean rank 4
Median rank 4
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name claude-fable-5
Benchmark version live leaderboard
Source rank 4
Rank spread 1 / 7
Raw score 1504
Unit Elo
Confidence interval 1495 / 1513
Sample count 5792
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Pending evaluation GPT-6 Astra: pending_evaluation. New release — awaiting comparable task-level evidence. Claude Fable 5.1: pending_evaluation. New release — awaiting comparable task-level evidence. Compare this task on the human ranking → Terminal software engineering For repository work, terminal coding and software engineering.
Approximately tied at the top. Terminal-Bench 4.0 with the stated agent harness; does not measure standalone code reasoning. Single source family; independent cross-source confirmation is unavailable in this edition.
Category code
Metric / benchmark Terminal-Bench
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 GPT-6 Astra
Consensus #2 Claude Fable 5.1
Consensus #3 Claude Opus 5
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Terminal-Bench":1} Top 3 #1 GPT-6 Astra
Base model GPT-6 Astra
Provider OpenAI
Official model ID Not reported
Variant Codex / (max)
Reasoning effort max
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name GPT-6 Astra (max)
Benchmark version 4.0
Source rank 1
Rank spread Not reported / Not reported
Raw score 58.2
Unit percent
Confidence interval 55.400000000000006 / 61
Sample count Not reported
Preliminary false
Variant Codex / (max)
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.178Z
Evidence freshness unknown
Used in aggregation true #2 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Claude Code / (max)
Reasoning effort max
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Fable 5.1 (max)
Benchmark version 4.0
Source rank 2
Rank spread Not reported / Not reported
Raw score 57.9
Unit percent
Confidence interval 54.1 / 61.699999999999996
Sample count Not reported
Preliminary false
Variant Claude Code / (max)
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.178Z
Evidence freshness unknown
Used in aggregation true #3 Claude Opus 5
Base model Claude Opus 5
Provider Anthropic
Official model ID Not reported
Variant Claude Code / (max)
Reasoning effort max
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Opus 5 (max)
Benchmark version 4.0
Source rank 3
Rank spread Not reported / Not reported
Raw score 51.8
Unit percent
Confidence interval 48.4 / 55.199999999999996
Sample count Not reported
Preliminary false
Variant Claude Code / (max)
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.178Z
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Web development For building and editing interactive web interfaces.
Approximately tied at the top. WebDev overall preference. The leading models have overlapping rank spreads. Single source family; independent cross-source confirmation is unavailable in this edition.
Category code
Metric / benchmark Code Arena | WebDev🏆Overall
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 GPT-6 Astra
Consensus #2 Claude Fable 5.1
Consensus #3 Claude Opus 5
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1} Top 3 #1 GPT-6 Astra
Base model GPT-6 Astra
Provider OpenAI
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name gpt-6-astra-max
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 2
Raw score 1797
Unit Elo
Confidence interval 1773 / 1821
Sample count 1199
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true #2 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5.1-max
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 2
Raw score 1762
Unit Elo
Confidence interval 1746 / 1778
Sample count 2275
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true #3 Claude Opus 5
Base model Claude Opus 5
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-opus-5-max
Benchmark version live leaderboard
Source rank 3
Rank spread 3 / 6
Raw score 1688
Unit Elo
Confidence interval 1680 / 1696
Sample count 10904
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name claude-opus-5-high
Benchmark version live leaderboard
Source rank 7
Rank spread 5 / 7
Raw score 1661
Unit Elo
Confidence interval 1654 / 1668
Sample count 10928
Preliminary false
Variant High
Reasoning effort high
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation false Compare this task on the human ranking → AI agents For autonomous multi-step work with tools and external environments.
Approximately tied at the top. Agent Arena overall plus targeted clarification ability; no claim about every agent workload.
Category code
Metric / benchmark Agent Arena🏆Overall; HiL-Bench
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5.1
Consensus #2 Claude Opus 5
Consensus #3 Claude Fable 5
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 2
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":0.75,"Scale Labs":0.25} Top 3 #1 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name Claude Fable 5.1 (Max)
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 3
Raw score 15.87
Unit percent
Confidence interval 13.03 / 18.71
Sample count 6796
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Claude Fable 5.1
Benchmark version current leaderboard
Source rank 1
Rank spread Not reported / Not reported
Raw score 61.5
Unit percent
Confidence interval 55.03 / 67.97
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 Claude Opus 5
Base model Claude Opus 5
Provider Anthropic
Official model ID Not reported
Variant High
Reasoning effort high
Consensus rank 2
Evaluation status evaluated
Consensus score 0.72319732
Mean rank 1.5
Median rank 1.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name Claude Opus 5 (High)
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 4
Raw score 12.68
Unit percent
Confidence interval 10.95 / 14.41
Sample count 22594
Preliminary false
Variant High
Reasoning effort high
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Claude Opus 5 (Max)
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 6
Raw score 11.67
Unit percent
Confidence interval 9.71 / 13.629999999999999
Sample count 18112
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation false
Source model name Claude Opus 5
Benchmark version current leaderboard
Source rank 1
Rank spread Not reported / Not reported
Raw score 57
Unit percent
Confidence interval 51.519999999999996 / 62.480000000000004
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant High
Reasoning effort high
Consensus rank 3
Evaluation status evaluated
Consensus score 0.57300742
Mean rank 2.5
Median rank 2.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name Claude Fable 5 (High)
Benchmark version live leaderboard
Source rank 4
Rank spread 2 / 7
Raw score 10.23
Unit percent
Confidence interval 8.77 / 11.690000000000001
Sample count 36204
Preliminary false
Variant High
Reasoning effort high
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Claude Fable 5
Benchmark version current leaderboard
Source rank 1
Rank spread Not reported / Not reported
Raw score 56.33
Unit percent
Confidence interval 50.83 / 61.83
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Image understanding For reasoning about photos, diagrams and visual content.
Approximately tied at the top. Visual question-answering preference; does not establish performance on every vision subtask. Single source family; independent cross-source confirmation is unavailable in this edition.
Category image
Metric / benchmark Vision Arena🏆Overall
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5
Consensus #2 Claude Opus 4.7
Consensus #3 Qwen3.8 Max
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1} Top 3 #1 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 6
Raw score 1313
Unit Elo
Confidence interval 1305 / 1321
Sample count 10002
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-08-27
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true #2 Claude Opus 4.7
Base model Claude Opus 4.7
Provider Anthropic
Official model ID Not reported
Variant High
Reasoning effort high
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-opus-4-7-high
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 11
Raw score 1301
Unit Elo
Confidence interval 1294 / 1308
Sample count 21137
Preliminary false
Variant High
Reasoning effort high
Source date 2026-08-27
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name claude-opus-4-7
Benchmark version live leaderboard
Source rank 4
Rank spread 1 / 14
Raw score 1299
Unit Elo
Confidence interval 1292 / 1306
Sample count 21457
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-08-27
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation false #3 Qwen3.8 Max
Base model Qwen3.8 Max
Provider Alibaba
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name qwen3.8-max
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 15
Raw score 1300
Unit Elo
Confidence interval 1292 / 1308
Sample count 7244
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-08-27
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true Compare this task on the human ranking → Text from image (OCR) For reading text embedded in images and scanned content.
Approximately tied at the top. OCR task slice; includes confidence intervals and vote counts. Single source family; independent cross-source confirmation is unavailable in this edition.
Category image
Metric / benchmark Vision — OCR
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5
Consensus #2 Qwen3.8 Max
Consensus #3 Claude Opus 4.7
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1} Top 3 #1 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5
Benchmark version Not reported
Source rank 1
Rank spread 1 / 5
Raw score 1329
Unit Elo
Confidence interval 1320 / 1338
Sample count 7174
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:10:38.778Z
Evidence freshness fresh
Used in aggregation true #2 Qwen3.8 Max
Base model Qwen3.8 Max
Provider Alibaba
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name qwen3.8-max
Benchmark version Not reported
Source rank 2
Rank spread 1 / 13
Raw score 1315
Unit Elo
Confidence interval 1306 / 1324
Sample count 5092
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:10:38.778Z
Evidence freshness fresh
Used in aggregation true #3 Claude Opus 4.7
Base model Claude Opus 4.7
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-opus-4-7
Benchmark version Not reported
Source rank 3
Rank spread 1 / 12
Raw score 1314
Unit Elo
Confidence interval 1307 / 1321
Sample count 15512
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:10:38.778Z
Evidence freshness fresh
Used in aggregation true
Source model name claude-opus-4-7-high
Benchmark version Not reported
Source rank 4
Rank spread 1 / 12
Raw score 1314
Unit Elo
Confidence interval 1307 / 1321
Sample count 15236
Preliminary false
Variant High
Reasoning effort high
Source date 2026-09-02
Captured at 2026-09-07T10:10:38.778Z
Evidence freshness fresh
Used in aggregation false Compare this task on the human ranking → Image generation For generating images from natural-language prompts.
Independent image preference sources; configurations are retained separately.
Category image
Metric / benchmark Text-to-Image Arena🏆Overall; Text to Image Leaderboard - Top AI Image Models | Artificial Analysis
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 GPT Image 2
Consensus #2 MAI-Image 2.6
Consensus #3 Reve 2.1
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 2
Source coverage 1
Confidence medium
Disagreement status clear_winner
CI overlap false
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1,"Artificial Analysis":1} Top 3 #1 GPT Image 2
Base model GPT Image 2
Provider OpenAI
Official model ID Not reported
Variant Medium
Reasoning effort medium
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name gpt-image-2 (medium)
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 1
Raw score 1382
Unit Elo
Confidence interval 1378 / 1386
Sample count 77830
Preliminary false
Variant Medium
Reasoning effort medium
Source date 2026-09-04
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name GPT Image 2 (high)
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 1
Raw score 1178
Unit Elo
Confidence interval 1168 / 1188
Sample count 14585
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 MAI-Image 2.6
Base model MAI-Image 2.6
Provider Microsoft AI
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name mai-image-2.6
Benchmark version live leaderboard
Source rank 2
Rank spread 2 / 3
Raw score 1332
Unit Elo
Confidence interval 1325 / 1339
Sample count 10086
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-04
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name MAI-Image-2.6
Benchmark version live leaderboard
Source rank 2
Rank spread 2 / 2
Raw score 1149
Unit Elo
Confidence interval 1137 / 1161
Sample count 6351
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 Reve 2.1
Base model Reve 2.1
Provider Reve
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.46533828
Mean rank 3.5
Median rank 3.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name reve-2.1
Benchmark version live leaderboard
Source rank 4
Rank spread 3 / 4
Raw score 1301
Unit Elo
Confidence interval 1293 / 1309
Sample count 7576
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-04
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Reve 2.1
Benchmark version live leaderboard
Source rank 3
Rank spread 3 / 4
Raw score 1127
Unit Elo
Confidence interval 1118 / 1136
Sample count 16045
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Image editing For editing existing images from text instructions.
Sources disagree on #1. Both GPT Image 2 and MAI-Image 2.6 are top tier.
Category image
Metric / benchmark Image Edit Arena🏆Single Image Edit; Image Editing Leaderboard - Top AI Image Models | Artificial Analysis
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 GPT Image 2
Consensus #2 MAI-Image 2.6
Consensus #3 Muse Image
Consensus score 0.81546488
Mean rank 1.5
Median rank 1.5
Independent source count 2
Source coverage 1
Confidence medium
Disagreement status source_split
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1,"Artificial Analysis":1} Top 3 #1 GPT Image 2
Base model GPT Image 2
Provider OpenAI
Official model ID Not reported
Variant Medium
Reasoning effort medium
Consensus rank 1
Evaluation status evaluated
Consensus score 0.81546488
Mean rank 1.5
Median rank 1.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name gpt-image-2 (medium)
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 1
Raw score 1461
Unit Elo
Confidence interval 1457 / 1465
Sample count 228827
Preliminary false
Variant Medium
Reasoning effort medium
Source date 2026-09-03
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name GPT Image 2 (high)
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 2
Raw score 1117
Unit Elo
Confidence interval 1107 / 1127
Sample count 11046
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 MAI-Image 2.6
Base model MAI-Image 2.6
Provider Microsoft AI
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.81546488
Mean rank 1.5
Median rank 1.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name mai-image-2.6
Benchmark version live leaderboard
Source rank 2
Rank spread 2 / 3
Raw score 1439
Unit Elo
Confidence interval 1429 / 1449
Sample count 5086
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-03
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name MAI-Image-2.6
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 2
Raw score 1122
Unit Elo
Confidence interval 1112 / 1132
Sample count 11173
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 Muse Image
Base model Muse Image
Provider Meta
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.46533828
Mean rank 3.5
Median rank 3.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name muse-image
Benchmark version live leaderboard
Source rank 4
Rank spread 4 / 5
Raw score 1403
Unit Elo
Confidence interval 1398 / 1408
Sample count 76663
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-03
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Muse Image
Benchmark version live leaderboard
Source rank 3
Rank spread 3 / 6
Raw score 1105
Unit Elo
Confidence interval 1095 / 1115
Sample count 5493
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Text → video For generating complete videos from text prompts.
Sources disagree on #1. Independent preference sources disagree; overlapping intervals prevent a definitive winner.
Category video
Metric / benchmark Text-to-Video Arena; Text to Video Leaderboard - Top AI Video Models | Artificial Analysis
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Wan 3.0
Consensus #2 Gemini Omni Flash
Consensus #3 MiniMax H3
Consensus score 0.75
Mean rank 2
Median rank 2
Independent source count 2
Source coverage 1
Confidence medium
Disagreement status source_split
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1,"Artificial Analysis":1} Top 3 #1 Wan 3.0
Base model Wan 3.0
Provider Alibaba
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 0.75
Mean rank 2
Median rank 2
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name wan3.0
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 7
Raw score 1494
Unit Elo
Confidence interval 1475 / 1513
Sample count 1167
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-04
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Wan 3.0
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 3
Raw score 1239
Unit Elo
Confidence interval 1229 / 1249
Sample count 5591
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 Gemini Omni Flash
Base model Gemini Omni Flash
Provider Google
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name gemini-omni-flash
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 5
Raw score 1511
Unit Elo
Confidence interval 1501 / 1521
Sample count 21934
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-04
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Gemini Omni Flash
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 3
Raw score 1238
Unit Elo
Confidence interval 1232 / 1244
Sample count 18134
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 MiniMax H3
Base model MiniMax H3
Provider MiniMax
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.40773244
Mean rank 5.5
Median rank 5.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name minimax-h3
Benchmark version live leaderboard
Source rank 8
Rank spread 6 / 9
Raw score 1462
Unit Elo
Confidence interval 1452 / 1472
Sample count 7648
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-04
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Minimax H3 Max (post-trained by fal)
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 4
Raw score 1235
Unit Elo
Confidence interval 1225 / 1245
Sample count 5451
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true
Source model name MiniMax H3
Benchmark version live leaderboard
Source rank 4
Rank spread 4 / 5
Raw score 1227
Unit Elo
Confidence interval 1220 / 1234
Sample count 8993
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation false Compare this task on the human ranking → Image → video For animating a still image into a generated video.
Sources disagree on ranking order. H3 Max is a fal post-trained configuration. Source-specific configuration evidence is not interchangeable.
Category video
Metric / benchmark Image-to-Video Arena; Image to Video Leaderboard - Top AI Video Models | Artificial Analysis
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 MiniMax H3
Consensus #2 Dreamina Seedance 2.0
Consensus #3 Wan 3.0
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 2
Source coverage 1
Confidence medium
Disagreement status source_split
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1,"Artificial Analysis":1} Top 3 #1 MiniMax H3
Base model MiniMax H3
Provider MiniMax
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name minimax-h3
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 3
Raw score 1497
Unit Elo
Confidence interval 1491 / 1503
Sample count 36137
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Minimax H3 Max (post-trained by fal)
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 2
Raw score 1201
Unit Elo
Confidence interval 1192 / 1210
Sample count 5587
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true
Source model name MiniMax H3
Benchmark version live leaderboard
Source rank 3
Rank spread 2 / 4
Raw score 1187
Unit Elo
Confidence interval 1179 / 1195
Sample count 7465
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation false #2 Dreamina Seedance 2.0
Base model Dreamina Seedance 2.0
Provider Bytedance
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.50889128
Mean rank 3.5
Median rank 3.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name dreamina-seedance-2.0-720p
Benchmark version live leaderboard
Source rank 5
Rank spread 2 / 5
Raw score 1477
Unit Elo
Confidence interval 1469 / 1485
Sample count 113611
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Dreamina Seedance 2.0 720p
Benchmark version live leaderboard
Source rank 2
Rank spread 2 / 3
Raw score 1192
Unit Elo
Confidence interval 1186 / 1198
Sample count 18094
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 Wan 3.0
Base model Wan 3.0
Provider Alibaba
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.4434264
Mean rank 4
Median rank 4
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name wan3.0
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 6
Raw score 1481
Unit Elo
Confidence interval 1465 / 1497
Sample count 1401
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-09-02
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Wan 3.0
Benchmark version live leaderboard
Source rank 5
Rank spread 4 / 5
Raw score 1176
Unit Elo
Confidence interval 1168 / 1184
Sample count 7791
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Video editing For instruction-driven edits and transformations of existing video.
Approximately tied at the top. Both sources put Wan 3.0 first, with uncertainty on Arena.
Category video
Metric / benchmark Video Edit Arena; Video Editing Leaderboard - Top AI Video Models | Artificial Analysis
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Wan 3.0
Consensus #2 MiniMax H3
Consensus #3 Gemini Omni Flash
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 2
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1,"Artificial Analysis":1} Top 3 #1 Wan 3.0
Base model Wan 3.0
Provider Alibaba
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name wan3.0
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 3
Raw score 1414
Unit Elo
Confidence interval 1388 / 1440
Sample count 463
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-08-26
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Wan 3.0
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 1
Raw score 1190
Unit Elo
Confidence interval 1183 / 1197
Sample count 5315
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 MiniMax H3
Base model MiniMax H3
Provider MiniMax
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.56546488
Mean rank 2.5
Median rank 2.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name minimax-h3
Benchmark version live leaderboard
Source rank 3
Rank spread 1 / 5
Raw score 1392
Unit Elo
Confidence interval 1373 / 1411
Sample count 962
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-08-26
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name MiniMax H3
Benchmark version live leaderboard
Source rank 2
Rank spread 2 / 3
Raw score 1129
Unit Elo
Confidence interval 1124 / 1134
Sample count 13014
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 Gemini Omni Flash
Base model Gemini Omni Flash
Provider Google
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.46533828
Mean rank 3.5
Median rank 3.5
Independent source count 2
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name gemini-omni-flash
Benchmark version live leaderboard
Source rank 4
Rank spread 3 / 5
Raw score 1367
Unit Elo
Confidence interval 1352 / 1382
Sample count 2109
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date 2026-08-26
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true
Source model name Gemini Omni Flash
Benchmark version live leaderboard
Source rank 3
Rank spread 2 / 3
Raw score 1125
Unit Elo
Confidence interval 1120 / 1130
Sample count 15995
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Speech to text For transcribing spoken audio into text.
AA-WER transcription accuracy only. Lower WER is better; latency is a separate criterion. Single source family; independent cross-source confirmation is unavailable in this edition.
Category audio
Metric / benchmark AA-WER Non-streaming
Consensus metric direction higher_is_better
Raw metric direction lower_is_better
Consensus #1 Fun-Realtime-ASR-preview
Consensus #2 MAI-Transcribe-2
Consensus #3 ElevenLabs Scribe v2
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status clear_winner
CI overlap false
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Artificial Analysis":1} Top 3 #1 Fun-Realtime-ASR-preview
Base model Fun-Realtime-ASR-preview
Provider Alibaba
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Fun-Realtime-ASR-preview
Benchmark version Not reported
Source rank 1
Rank spread Not reported / Not reported
Raw score 1.7
Unit percent WER
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 MAI-Transcribe-2
Base model MAI-Transcribe-2
Provider Microsoft AI
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name MAI-Transcribe-2
Benchmark version Not reported
Source rank 2
Rank spread Not reported / Not reported
Raw score 2
Unit percent WER
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 ElevenLabs Scribe v2
Base model ElevenLabs Scribe v2
Provider ElevenLabs
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name ElevenLabs Scribe v2
Benchmark version Not reported
Source rank 3
Rank spread Not reported / Not reported
Raw score 2.2
Unit percent WER
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Text to speech For generating natural spoken audio from text.
Approximately tied at the top. Provider Voice Arena measures preference with each provider’s selected voices. Single source family; independent cross-source confirmation is unavailable in this edition.
Category audio
Metric / benchmark Text to Speech Leaderboard - Top AI Speech Models | Artificial Analysis
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Sonic 3.6
Consensus #2 Realtime TTS-2
Consensus #3 Qwen-Audio-3.0-TTS-Plus
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Artificial Analysis":1} Top 3 #1 Sonic 3.6
Base model Sonic 3.6
Provider Cartesia
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Sonic 3.6
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 1
Raw score 1282
Unit Elo
Confidence interval 1265 / 1299
Sample count 1790
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 Realtime TTS-2
Base model Realtime TTS-2
Provider Inworld
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Realtime TTS-2
Benchmark version live leaderboard
Source rank 2
Rank spread 2 / 4
Raw score 1252
Unit Elo
Confidence interval 1234 / 1270
Sample count 1094
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 Qwen-Audio-3.0-TTS-Plus
Base model Qwen-Audio-3.0-TTS-Plus
Provider Alibaba
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Qwen-Audio-3.0-TTS-Plus
Benchmark version live leaderboard
Source rank 3
Rank spread 2 / 5
Raw score 1241
Unit Elo
Confidence interval 1227 / 1255
Sample count 2398
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Voice agents For voice agents completing customer-service tasks through spoken interaction.
Voice agents ranked by customer-service task completion (tau-Voice); conversation is a separate task. Single source family; independent cross-source confirmation is unavailable in this edition.
Category audio
Metric / benchmark Voice Agentic Performance
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Grok Voice Think Fast 2.0
Consensus #2 Qwen Audio 3.0 Realtime Plus
Consensus #3 Grok Voice Think Fast 1.0
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status clear_winner
CI overlap false
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Artificial Analysis":1} Top 3 #1 Grok Voice Think Fast 2.0
Base model Grok Voice Think Fast 2.0
Provider SpaceXAI
Official model ID Not reported
Variant High
Reasoning effort high
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Grok Voice Think Fast 2.0 High
Benchmark version current components
Source rank 1
Rank spread Not reported / Not reported
Raw score 56.5
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation true
Speech to Speech Index % 79.0%
Speech Reasoning % 97%
Conversational Dynamics % 95.1%
Agentic Performance % 56.5%
Arena Preference Elo 908
Task Success Rate % 94.7%
Time to First Audio Seconds 0.70 #2 Qwen Audio 3.0 Realtime Plus
Base model Qwen Audio 3.0 Realtime Plus
Provider Alibaba Cloud
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Qwen Audio 3.0 Realtime Plus
Benchmark version current components
Source rank 2
Rank spread Not reported / Not reported
Raw score 54.6
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation true
Speech to Speech Index % 66.8%
Speech Reasoning % 99%
Conversational Dynamics % 98.4%
Agentic Performance % 54.6%
Arena Preference Elo 699
Task Success Rate % 77.8%
Time to First Audio Seconds 1.54 #3 Grok Voice Think Fast 1.0
Base model Grok Voice Think Fast 1.0
Provider SpaceXAI
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Grok Voice Think Fast 1.0
Benchmark version current components
Source rank 3
Rank spread Not reported / Not reported
Raw score 52.1
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation true
Speech to Speech Index % 72.3%
Speech Reasoning % 97%
Conversational Dynamics % 77.8%
Agentic Performance % 52.1%
Arena Preference Elo 839
Task Success Rate % 80.7%
Time to First Audio Seconds 1.25 Compare this task on the human ranking → Music generation For generating complete instrumental or vocal music.
Approximately tied at the top. Instrumental and vocal preference remain separate lists. Single source family; independent cross-source confirmation is unavailable in this edition.
Category audio
Metric / benchmark Music Arena — Instrumental
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Suno V5.5
Consensus #2 Mureka V9
Consensus #3 Mureka V8
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Artificial Analysis":1} Instrumental #1 Suno V5.5
Base model Suno V5.5
Provider Suno
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Suno V5.5
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 1
Raw score 1186
Unit Elo
Confidence interval 1177 / 1195
Sample count 6469
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:29.413Z
Evidence freshness unknown
Used in aggregation true #2 Mureka V9
Base model Mureka V9
Provider Mureka
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Mureka V9
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 3
Raw score 1177
Unit Elo
Confidence interval 1163 / 1191
Sample count 3046
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:29.413Z
Evidence freshness unknown
Used in aggregation true #3 Mureka V8
Base model Mureka V8
Provider Mureka
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Mureka V8
Benchmark version live leaderboard
Source rank 3
Rank spread 3 / 4
Raw score 1166
Unit Elo
Confidence interval 1158 / 1174
Sample count 7497
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:29.413Z
Evidence freshness unknown
Used in aggregation true Vocals #1 Suno V5.5
Base model Suno V5.5
Provider Suno
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Suno V5.5
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 1
Raw score 1170
Unit Elo
Confidence interval 1163 / 1177
Sample count 10373
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:29.205Z
Evidence freshness unknown
Used in aggregation true #2 Mureka V9
Base model Mureka V9
Provider Mureka
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Mureka V9
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 2
Raw score 1159
Unit Elo
Confidence interval 1145 / 1173
Sample count 3010
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:29.205Z
Evidence freshness unknown
Used in aggregation true #3 Mureka V8
Base model Mureka V8
Provider Mureka
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Mureka V8
Benchmark version live leaderboard
Source rank 3
Rank spread 3 / 3
Raw score 1140
Unit Elo
Confidence interval 1134 / 1146
Sample count 11414
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:29.205Z
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Data analysis For objective data-analysis problems; dashboard design and standalone SQL quality are separate tasks.
Objective data analysis; this is not a web dashboard design benchmark and does not establish a separate SQL winner. Single source family; independent cross-source confirmation is unavailable in this edition.
Category code
Metric / benchmark LiveBench Data Analysis
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 GPT-6 Astra
Consensus #2 GPT-5.5
Consensus #3 Claude Fable 5
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status clear_winner
CI overlap false
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"LiveBench":1} Top 3 #1 GPT-6 Astra
Base model GPT-6 Astra
Provider OpenAI
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name GPT-6 Astra Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 1
Rank spread Not reported / Not reported
Raw score 83
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true #2 GPT-5.5
Base model GPT-5.5
Provider OpenAI
Official model ID Not reported
Variant Xhigh
Reasoning effort xhigh
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name GPT-5.5 Thinking xHigh Effort
Benchmark version 2026-06-25 (latest release)
Source rank 2
Rank spread Not reported / Not reported
Raw score 81.6
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Xhigh
Reasoning effort xhigh
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true #3 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Claude Fable 5 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 3
Rank spread Not reported / Not reported
Raw score 80.5
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Structured output and tool calling For reliable function calls, JSON arguments and structured tool use.
BFCL function-calling accuracy, with explicit benchmark configurations. Single source family; independent cross-source confirmation is unavailable in this edition. Warning: the published benchmark snapshot is more than 30 days old.
Category code
Metric / benchmark BFCL overall function calling
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude-Opus-4.5-20251101
Consensus #2 Claude-Sonnet-4.5-20250929
Consensus #3 Gemini-3-Pro-Preview
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence low
Disagreement status clear_winner
CI overlap false
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Berkeley BFCL":1} Top 3 #1 Claude-Opus-4.5-20251101
Base model Claude-Opus-4.5-20251101
Provider Anthropic
Official model ID Not reported
Variant function calling
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence low Source-specific results
Source model name Claude-Opus-4-5-20251101 (FC)
Benchmark version V4 / f7cf735
Source rank 1
Rank spread Not reported / Not reported
Raw score 77.47
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant function calling
Reasoning effort Not reported
Source date 2026-04-12
Captured at 2026-09-07T10:10:33.001Z
Evidence freshness stale
Used in aggregation true
Source model name Claude-Opus-4-5-20251101 (Prompt)
Benchmark version V4 / f7cf735
Source rank 57
Rank spread Not reported / Not reported
Raw score 33.47
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant function calling
Reasoning effort Not reported
Source date 2026-04-12
Captured at 2026-09-07T10:10:33.001Z
Evidence freshness stale
Used in aggregation false #2 Claude-Sonnet-4.5-20250929
Base model Claude-Sonnet-4.5-20250929
Provider Anthropic
Official model ID Not reported
Variant function calling
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence low Source-specific results
Source model name Claude-Sonnet-4-5-20250929 (FC)
Benchmark version V4 / f7cf735
Source rank 2
Rank spread Not reported / Not reported
Raw score 73.24
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant function calling
Reasoning effort Not reported
Source date 2026-04-12
Captured at 2026-09-07T10:10:33.001Z
Evidence freshness stale
Used in aggregation true
Source model name Claude-Sonnet-4-5-20250929 (Prompt)
Benchmark version V4 / f7cf735
Source rank 89
Rank spread Not reported / Not reported
Raw score 24.9
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant function calling
Reasoning effort Not reported
Source date 2026-04-12
Captured at 2026-09-07T10:10:33.001Z
Evidence freshness stale
Used in aggregation false #3 Gemini-3-Pro-Preview
Base model Gemini-3-Pro-Preview
Provider Google
Official model ID Not reported
Variant function calling
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence low Source-specific results
Source model name Gemini-3-Pro-Preview (Prompt)
Benchmark version V4 / f7cf735
Source rank 3
Rank spread Not reported / Not reported
Raw score 72.51
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant function calling
Reasoning effort Not reported
Source date 2026-04-12
Captured at 2026-09-07T10:10:33.001Z
Evidence freshness stale
Used in aggregation true
Source model name Gemini-3-Pro-Preview (FC)
Benchmark version V4 / f7cf735
Source rank 7
Rank spread Not reported / Not reported
Raw score 68.14
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant function calling
Reasoning effort Not reported
Source date 2026-04-12
Captured at 2026-09-07T10:10:33.001Z
Evidence freshness stale
Used in aggregation false Compare this task on the human ranking → Controlled-voice synthesis For comparing speech synthesis with a controlled common voice; speaker similarity is not evaluated.
Approximately tied at the top. Controlled Voice Arena tests common-voice synthesis preference; it does not separately establish speaker identity similarity. Single source family; independent cross-source confirmation is unavailable in this edition.
Category audio
Metric / benchmark Controlled Voice Arena
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Realtime TTS-2
Consensus #2 Sonic 3.6
Consensus #3 Sonic 3.5
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence low
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Artificial Analysis":1} Top 3 #1 Realtime TTS-2
Base model Realtime TTS-2
Provider Inworld
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence low Source-specific results
Source model name Realtime TTS-2
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 2
Raw score 1123
Unit Elo
Confidence interval 1107 / 1139
Sample count 1292
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:32.121Z
Evidence freshness unknown
Used in aggregation true #2 Sonic 3.6
Base model Sonic 3.6
Provider Cartesia
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence low Source-specific results
Source model name Sonic 3.6
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 2
Raw score 1119
Unit Elo
Confidence interval 1104 / 1134
Sample count 1993
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:32.121Z
Evidence freshness unknown
Used in aggregation true #3 Sonic 3.5
Base model Sonic 3.5
Provider Cartesia
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence low Source-specific results
Source model name Sonic 3.5
Benchmark version live leaderboard
Source rank 3
Rank spread 3 / 3
Raw score 1096
Unit Elo
Confidence interval 1084 / 1108
Sample count 3791
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:32.121Z
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Code reasoning Objective coding problems without a terminal agent harness.
Objective coding problems without a terminal agent harness. Single source family; independent cross-source confirmation is unavailable in this edition.
Category code
Metric / benchmark LiveBench Coding
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5.1
Consensus #2 Claude Fable 5
Consensus #3 GPT-5.6 Sol
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status clear_winner
CI overlap false
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"LiveBench":1} Top 3 #1 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Claude Fable 5.1 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 1
Rank spread Not reported / Not reported
Raw score 86.4
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true #2 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Claude Fable 5 Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 2
Rank spread Not reported / Not reported
Raw score 86
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true #3 GPT-5.6 Sol
Base model GPT-5.6 Sol
Provider OpenAI
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name GPT-5.6 Sol Max Effort
Benchmark version 2026-06-25 (latest release)
Source rank 3
Rank spread Not reported / Not reported
Raw score 83.9
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Max
Reasoning effort max
Source date Not reported
Captured at 2026-09-07T10:10:19.053Z
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Reference-based frontend design Recreate a frontend from a design reference; separate from general WebDev.
Approximately tied at the top. Recreate a frontend from a design reference; separate from general WebDev. Single source family; independent cross-source confirmation is unavailable in this edition.
Category code
Metric / benchmark Code Arena | WebDev🎨Reference-Based Design
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5.1
Consensus #2 GPT-6 Astra
Consensus #3 Kimi K3 Max
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Arena":1} Top 3 #1 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name claude-fable-5.1-max
Benchmark version live leaderboard
Source rank 1
Rank spread 1 / 2
Raw score 1815
Unit Elo
Confidence interval 1773 / 1857
Sample count 366
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true #2 GPT-6 Astra
Base model GPT-6 Astra
Provider OpenAI
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 2
Evaluation status evaluated
Consensus score 0.63092975
Mean rank 2
Median rank 2
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name gpt-6-astra-max
Benchmark version live leaderboard
Source rank 2
Rank spread 1 / 3
Raw score 1785
Unit Elo
Confidence interval 1731 / 1839
Sample count 220
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true #3 Kimi K3 Max
Base model Kimi K3 Max
Provider Moonshot
Official model ID Not reported
Variant Max
Reasoning effort max
Consensus rank 3
Evaluation status evaluated
Consensus score 0.5
Mean rank 3
Median rank 3
Independent source count 1
Source coverage 1
Fresh source count 1
Confidence medium Source-specific results
Source model name kimi-k3-max
Benchmark version live leaderboard
Source rank 3
Rank spread 2 / 8
Raw score 1723
Unit Elo
Confidence interval 1695 / 1751
Sample count 777
Preliminary false
Variant Max
Reasoning effort max
Source date 2026-09-05
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness fresh
Used in aggregation true Compare this task on the human ranking → Repository refactoring Restructure repositories while preserving behavior, using the stated agent harness.
Approximately tied at the top. Restructure repositories while preserving behavior, using the stated agent harness. Single source family; independent cross-source confirmation is unavailable in this edition.
Category code
Metric / benchmark SWE Atlas — Refactoring
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Claude Fable 5.1
Consensus #2 Claude Fable 5
Consensus #3 Claude Opus 4.7
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status statistical_tie
CI overlap true
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Scale Labs":1} Top 3 #1 Claude Fable 5.1
Base model Claude Fable 5.1
Provider Anthropic
Official model ID Not reported
Variant Xhigh
Reasoning effort xhigh
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Fable-5.1 (Claude Code) xHigh
Benchmark version current leaderboard
Source rank 1
Rank spread Not reported / Not reported
Raw score 56.67
Unit percent
Confidence interval 50.150000000000006 / 63.19
Sample count Not reported
Preliminary false
Variant Xhigh
Reasoning effort xhigh
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #2 Claude Fable 5
Base model Claude Fable 5
Provider Anthropic
Official model ID Not reported
Variant Xhigh
Reasoning effort xhigh
Consensus rank 2
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Fable-5 (Claude Code) xHigh
Benchmark version current leaderboard
Source rank 1
Rank spread Not reported / Not reported
Raw score 54.76
Unit percent
Confidence interval 48 / 61.519999999999996
Sample count Not reported
Preliminary false
Variant Xhigh
Reasoning effort xhigh
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true #3 Claude Opus 4.7
Base model Claude Opus 4.7
Provider Anthropic
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 3
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Opus-4.7 (Claude Code)
Benchmark version current leaderboard
Source rank 1
Rank spread Not reported / Not reported
Raw score 48.57
Unit percent
Confidence interval 41.84 / 55.3
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:16:44.291125+00:00
Evidence freshness unknown
Used in aggregation true Compare this task on the human ranking → Realtime voice conversation Equal rank weight for speech reasoning and conversational dynamics. The composite index and agent completion are retained as component context.
Sources disagree on ranking order. Equal rank weight for speech reasoning and conversational dynamics. The composite index and agent completion are retained as component context. Single source family; independent cross-source confirmation is unavailable in this edition.
Category audio
Metric / benchmark Speech Reasoning; Conversational Dynamics
Consensus metric direction higher_is_better
Raw metric direction higher_is_better
Consensus #1 Qwen Audio 3.0 Realtime Plus
Consensus #2 Qwen Audio 3.0 Realtime Flash
Consensus #3 GPT-Realtime-2
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Confidence medium
Disagreement status source_split
CI overlap false
Last verified 2026-09-07T10:16:44.291125+00:00
Source family weights {"Artificial Analysis":1} Top 3 #1 Qwen Audio 3.0 Realtime Plus
Base model Qwen Audio 3.0 Realtime Plus
Provider Alibaba Cloud
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 1
Evaluation status evaluated
Consensus score 1
Mean rank 1
Median rank 1
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Qwen Audio 3.0 Realtime Plus
Benchmark version current components
Source rank 1
Rank spread Not reported / Not reported
Raw score 99
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation true
Speech to Speech Index % 66.8%
Speech Reasoning % 99%
Conversational Dynamics % 98.4%
Agentic Performance % 54.6%
Arena Preference Elo 699
Task Success Rate % 77.8%
Time to First Audio Seconds 1.54
Source model name Qwen Audio 3.0 Realtime Plus
Benchmark version current components
Source rank 1
Rank spread Not reported / Not reported
Raw score 98.4
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation true
Speech to Speech Index % 66.8%
Speech Reasoning % 99%
Conversational Dynamics % 98.4%
Agentic Performance % 54.6%
Arena Preference Elo 699
Task Success Rate % 77.8%
Time to First Audio Seconds 1.54 #2 Qwen Audio 3.0 Realtime Flash
Base model Qwen Audio 3.0 Realtime Flash
Provider Alibaba Cloud
Official model ID Not reported
Variant Not reported
Reasoning effort Not reported
Consensus rank 2
Evaluation status evaluated
Consensus score 0.47319732
Mean rank 5
Median rank 5
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name Qwen Audio 3.0 Realtime Flash
Benchmark version current components
Source rank 8
Rank spread Not reported / Not reported
Raw score 96
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation true
Speech to Speech Index % 64.2%
Speech Reasoning % 96%
Conversational Dynamics % 96.9%
Agentic Performance % 35.9%
Arena Preference Elo 752
Task Success Rate % 81.7%
Time to First Audio Seconds 1.55
Source model name Qwen Audio 3.0 Realtime Flash
Benchmark version current components
Source rank 2
Rank spread Not reported / Not reported
Raw score 96.9
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Not reported
Reasoning effort Not reported
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation true
Speech to Speech Index % 64.2%
Speech Reasoning % 96%
Conversational Dynamics % 96.9%
Agentic Performance % 35.9%
Arena Preference Elo 752
Task Success Rate % 81.7%
Time to First Audio Seconds 1.55 #3 GPT-Realtime-2
Base model GPT-Realtime-2
Provider OpenAI
Official model ID Not reported
Variant High
Reasoning effort high
Consensus rank 3
Evaluation status evaluated
Consensus score 0.46533828
Mean rank 3.5
Median rank 3.5
Independent source count 1
Source coverage 1
Fresh source count 0
Confidence medium Source-specific results
Source model name GPT-Realtime-2 (High)
Benchmark version current components
Source rank 4
Rank spread Not reported / Not reported
Raw score 97
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation true
Speech to Speech Index % 73.6%
Speech Reasoning % 97%
Conversational Dynamics % 95.3%
Agentic Performance % 39.8%
Arena Preference Elo 914
Task Success Rate % 89.8%
Time to First Audio Seconds 1.14
Source model name GPT-Realtime-2 (Medium)
Benchmark version current components
Source rank 10
Rank spread Not reported / Not reported
Raw score 93
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Medium
Reasoning effort medium
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation false
Speech to Speech Index % -
Speech Reasoning % 93%
Conversational Dynamics % 95.2%
Agentic Performance % 37.4%
Arena Preference Elo -
Task Success Rate % -
Time to First Audio Seconds 1.22
Source model name GPT-Realtime-2 (Minimal)
Benchmark version current components
Source rank 19
Rank spread Not reported / Not reported
Raw score 72
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Minimal
Reasoning effort minimal
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation false
Speech to Speech Index % 62.7%
Speech Reasoning % 72%
Conversational Dynamics % 96.1%
Agentic Performance % 30.8%
Arena Preference Elo 881
Task Success Rate % 84.7%
Time to First Audio Seconds 1.12
Source model name GPT-Realtime-2 (High)
Benchmark version current components
Source rank 7
Rank spread Not reported / Not reported
Raw score 95.3
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant High
Reasoning effort high
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation false
Speech to Speech Index % 73.6%
Speech Reasoning % 97%
Conversational Dynamics % 95.3%
Agentic Performance % 39.8%
Arena Preference Elo 914
Task Success Rate % 89.8%
Time to First Audio Seconds 1.14
Source model name GPT-Realtime-2 (Medium)
Benchmark version current components
Source rank 8
Rank spread Not reported / Not reported
Raw score 95.2
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Medium
Reasoning effort medium
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation false
Speech to Speech Index % -
Speech Reasoning % 93%
Conversational Dynamics % 95.2%
Agentic Performance % 37.4%
Arena Preference Elo -
Task Success Rate % -
Time to First Audio Seconds 1.22
Source model name GPT-Realtime-2 (Minimal)
Benchmark version current components
Source rank 3
Rank spread Not reported / Not reported
Raw score 96.1
Unit percent
Confidence interval Not reported / Not reported
Sample count Not reported
Preliminary false
Variant Minimal
Reasoning effort minimal
Source date Not reported
Captured at 2026-09-07T10:10:23.006Z
Evidence freshness unknown
Used in aggregation true
Speech to Speech Index % 62.7%
Speech Reasoning % 72%
Conversational Dynamics % 96.1%
Agentic Performance % 30.8%
Arena Preference Elo 881
Task Success Rate % 84.7%
Time to First Audio Seconds 1.12 Compare this task on the human ranking → Tasks awaiting comparable evidence Translation: unavailable. Previous ranking used an agency article and unnamed model families; comparable version-specific primary evidence has not been established.