LLL.bz / modelsRaw JSON

AI model ranking data for agents

Top models for all 26 LLL.bz tasks.

Start with the task, then select its winner or one of the two ranked alternatives. There is no universal best model, and scores from different benchmarks are not directly comparable.

Tasks
26
Task groups
5
Top-3 lists
27
Unique models
40

Default list per task

Current winners

For tasks with variants, this summary uses the first listed variant. Open the task before choosing.

  1. 1
    General chat

    Claude Fable 5Arena score 1508±5

    text
  2. 2
    Russian language

    Claude Fable 5Arena score 1518±13

    text
  3. 3
    Expert analysis

    Claude Fable 5Arena score 1548±12

    text
  4. 4
    Mathematics

    Claude Opus 5Arena score 1545±24

    text
  5. 5
    Instruction following

    Claude Opus 4.6Arena score 1514±5

    text
  6. 6
    Copywriting

    Claude Fable 5Arena score 1508±9

    text
  7. 7
    Search and research

    Claude Opus 4.6 SearchArena score 1253±5

    text
  8. 8
    PDF and documents

    Claude Opus 5Arena score 1520±15

    text
  9. 9
    Programming

    Claude Opus 5Terminal-Bench score 89.1%

    code
  10. 10
    Web development

    Claude Opus 5Arena score 1691±9

    code
  11. 11
    AI agents

    Claude Opus 5Net improvement 12.47%±1.54

    code
  12. 12
    Image understanding

    Claude Fable 5Arena score 1312±9

    image
  13. 13
    Text from image (OCR)

    Claude Fable 5Arena score 1329±9

    image
  14. 14
    Image generation

    GPT Image 2Elo 1369

    image
  15. 15
    Image editing

    GPT Image 2Arena score 1463±4

    image
  16. 16
    Text → video

    Wan 3.0Elo 1244

    video
  17. 17
    Image → video

    Gemini Omni FlashElo 1366

    video
  18. 18
    Video editing

    Wan 3.0Elo 1192

    video
  19. 19
    Speech to text

    Fun-Realtime-ASR-previewAA-WER 1.7%

    audio
  20. 20
    Text to speech

    Cartesia Sonic 3.6Elo 1285

    audio
  21. 21
    Voice agents

    Gemini 3.1 Flash Live PreviewElo 1046

    audio
  22. 22
    Music generation

    Suno V5.5Elo 1186

    audio
  23. 23
    Translation

    GeminiAQI 77.7

    text
  24. 24
    Data analysis and SQL

    Qwen3.8 MaxArena score 1614±30

    code
  25. 25
    Structured output and tool calling

    Qwen3.7 MaxBFCL V4 score 75.0%

    code
  26. 26
    Voice cloning

    Cartesia Sonic 3.6Rank 1

    audio

Task group · 9 tasks

Text

Open this group on LLL.bz

Task · Arena score

General chat

Open ranking

For broad conversational quality across everyday prompts.

Note: The primary ranking is a blind-preference arena; confidence intervals overlap.

Top 3

  1. #1
    Claude Fable 5

    Anthropic

    Ranks first in the current broad text preference arena.

    Arena score1508±5
  2. #2
    Claude Opus 4.6

    Anthropic · High

    A near-tied option for broad conversational prompts.

    Arena score1504±4
  3. #3
    Claude Opus 4.7

    Anthropic · High

    Places inside the same tightly clustered chat tier.

    Arena score1502±4

Task · Arena score

Russian language

Open ranking

For Russian-language conversation and writing.

Note: Fresh primary data exists, but three independent Russian-specific sources reproducing the full top three were not found.

Top 3

  1. #1
    Claude Fable 5

    Anthropic

    Leads the current Russian-language Arena slice.

    Arena score1518±13
  2. #2
    Gemini 3.7 Flash

    Google · High

    Scores effectively level with the current leader in Russian.

    Arena score1517±24
  3. #3
    Claude Opus 4.6

    Anthropic

    Remains inside the leading Russian preference cluster.

    Arena score1509±8

Task · Arena score

Expert analysis

Open ranking

For difficult analytical and professional knowledge tasks.

Note: The top three intervals overlap, so the exact order is not statistically decisive.

Top 3

  1. #1
    Claude Fable 5

    Anthropic

    Has the highest current score on expert-oriented Arena prompts.

    Arena score1548±12
  2. #2
    Claude Opus 4.6

    Anthropic · High

    Scores essentially level with first place on expert prompts.

    Arena score1546±9
  3. #3
    Claude Opus 5

    Anthropic · High

    Places in the leading cluster for specialist analysis.

    Arena score1541±12

Task · Arena score

Mathematics

Open ranking

For hard calculations, proofs and multi-step maths.

Note: All three reported confidence intervals overlap materially.

Top 3

  1. #1
    Claude Opus 5

    Anthropic · Max

    Has the highest displayed score in the current Math Arena slice.

    Arena score1545±24
  2. #2
    Claude Fable 5

    Anthropic

    Ranks inside the same leading hard-math cluster.

    Arena score1528±17
  3. #3
    Gemini 3.7 Flash

    Google · High

    Posts a near-tied preliminary score on difficult mathematics.

    Arena score1523±33

Task · Arena score

Instruction following

Open ranking

For prompts with detailed constraints and required formats.

Note: First and second are practically tied within the reported intervals.

Top 3

  1. #1
    Claude Opus 4.6

    Anthropic · High

    Edges the current instruction-following Arena by displayed score.

    Arena score1514±5
  2. #2
    Claude Fable 5

    Anthropic

    Scores effectively level with first place on constrained prompts.

    Arena score1512±7
  3. #3
    Claude Opus 4.7

    Anthropic · High

    Maintains a high preference score on detailed instructions.

    Arena score1503±6

Task · Arena score

Copywriting

Open ranking

For persuasive, creative and polished marketing text.

Note: The source measures blind writing preference, not measured conversion performance.

Top 3

  1. #1
    Claude Fable 5

    Anthropic

    Ranks first on current blind creative-writing preference.

    Arena score1508±9
  2. #2
    Claude Opus 4.6

    Anthropic · High

    Scores close to first place for creative prose.

    Arena score1499±7
  3. #3
    Gemini 3.7 Flash

    Google · High

    Places in the current leading creative-writing cluster.

    Arena score1493±18

Task · Arena score

Search and research

Open ranking

For web search, evidence gathering and sourced answers.

Note: The dedicated search board updates slower than the ten-day rebuild cycle.

Top 3

  1. #1
    Claude Opus 4.6 Search

    Anthropic

    Leads the dedicated search-and-retrieval board.

    Arena score1253±5
  2. #2
    GPT-5.5 Search

    OpenAI

    Places second on the same search board.

    Arena score1240±5
  3. #3
    Claude Fable 5

    Anthropic

    Completes the leading trio for search tasks.

    Arena score1237±8

Task · Arena score

PDF and documents

Open ranking

For understanding long documents, PDFs and mixed document content.

Note: The document board updates slower than the ten-day cycle; duplicate Claude Opus 4.6 configurations are collapsed into one entry.

Top 3

  1. #1
    Claude Opus 5

    Anthropic · High

    Leads the dedicated document-handling board.

    Arena score1520±15
  2. #2
    Claude Opus 4.6

    Anthropic

    Ranks second after collapsing duplicate configurations.

    Arena score1510±6
  3. #3
    Claude Fable 5

    Anthropic

    Becomes third after enforcing one configuration per base model.

    Arena score1504±9

Task · AQI

Translation

Open ranking

For translating text accurately between languages.

Note: Measured by model family rather than exact version: the evaluation reports Gemini, Claude and GPT as families, so read this as a family-level ranking.

Top 3

  1. #1
    Gemini

    Google

    Ranks first for translation quality across the tested sample.

    AQI77.7
  2. #2
    Claude

    Anthropic

    Ranks second across the same translation sample.

    AQI75.6
  3. #3
    GPT

    OpenAI

    Ranks third across the same translation sample.

    AQI73.1

Task group · 5 tasks

Code

Open this group on LLL.bz

Task · Terminal-Bench score

Programming

Open ranking

For repository work, terminal coding and software engineering.

Note: Terminal-Bench measures agentic terminal coding; SWE-bench and Code Arena measure different coding workloads.

Top 3

  1. #1
    Claude Opus 5

    Anthropic · Max

    Has the highest reported Terminal-Bench result of the three.

    Terminal-Bench score89.1%
  2. #2
    GPT-5.6 Sol

    OpenAI

    Trails the reported leader by three tenths of a point.

    Terminal-Bench score88.8%
  3. #3
    Grok 4.6

    xAI

    Places inside the same tightly grouped Terminal-Bench frontier.

    Terminal-Bench score88.4%

Task · Arena score

Web development

Open ranking

For building and editing interactive web interfaces.

Note: The ranking measures blind preference over generated web applications.

Top 3

  1. #1
    Claude Opus 5

    Anthropic · Max

    Leads the current WebDev Arena by displayed score.

    Arena score1691±9
  2. #2
    Kimi K3 Max

    Moonshot AI

    Places second on blind evaluation of generated web applications.

    Arena score1674±11
  3. #3
    Qwen3.8 Max

    Alibaba

    Places close behind second in the current WebDev Arena.

    Arena score1669±13

Task · Net improvement

AI agents

Open ranking

For autonomous multi-step work with tools and external environments.

Metric: Higher net improvement is better.

Note: Claude Opus 5 Max is omitted because it is another configuration of the same base model.

Top 3

  1. #1
    Claude Opus 5

    Anthropic · High

    Has the highest distinct-base-model result in the current Agent Arena.

    Net improvement12.47%±1.54
  2. #2
    Claude Fable 5

    Anthropic · High

    Ranks second after collapsing duplicate Opus 5 configurations.

    Net improvement11.57%±1.70
  3. #3
    Kimi K3 Max

    Moonshot AI

    Is the next distinct base model in the Agent Arena ranking.

    Net improvement10.41%±0.62

Task · Arena score

Data analysis and SQL

Open ranking

For tables, CSV analysis, SQL generation and data-driven conclusions.

Note: The fresh source measures data-oriented web applications, not normalized text-to-SQL execution accuracy. BIRD and Spider2 did not yield a clean current raw-model top three that could be reconciled with this board.

Top 3

  1. #1
    Qwen3.8 Max

    Alibaba

    Leads the fresh Data & Analytics WebDev slice.

    Arena score1614±30
  2. #2
    Claude Opus 5

    Anthropic · Max

    Places effectively level with first on data-oriented web tasks.

    Arena score1612±16
  3. #3
    Kimi K3 Max

    Moonshot AI

    Completes the leading Data & Analytics WebDev cluster.

    Arena score1600±22

Task · BFCL V4 score

Structured output and tool calling

Open ranking

For reliable function calls, JSON arguments and structured tool use.

Note: Scores come from the Berkeley function-calling benchmark. Tool schemas differ between providers, so narrow gaps here mean little in practice.

Top 3

  1. #1
    Qwen3.7 Max

    Alibaba

    Has the highest BFCL V4 score of the three.

    BFCL V4 score75.0%
  2. #2
    BTL-4

    Bad Theory Labs

    Places second on the captured BFCL V4 score view.

    BFCL V4 score73.5%
  3. #3
    Ling 3.0 Flash

    InclusionAI

    Places third on the captured BFCL V4 score view.

    BFCL V4 score73.0%

Task group · 4 tasks

Images

Open this group on LLL.bz

Task · Arena score

Image understanding

Open ranking

For reasoning about photos, diagrams and visual content.

Note: Second and third are practically tied within the reported intervals.

Top 3

  1. #1
    Claude Fable 5

    Anthropic

    Leads the current blind-preference vision leaderboard.

    Arena score1312±9
  2. #2
    Qwen3.8 Max

    Alibaba

    Places narrowly ahead of third on mixed visual prompts.

    Arena score1302±8
  3. #3
    Claude Opus 4.7

    Anthropic · High

    Scores effectively level with second place on visual prompts.

    Arena score1301±7

Task · Arena score

Text from image (OCR)

Open ranking

For reading text embedded in images and scanned content.

Note: OCR Arena is preference-based rather than a character-error-rate benchmark.

Top 3

  1. #1
    Claude Fable 5

    Anthropic

    Leads the current OCR-specific Arena preference slice.

    Arena score1329±9
  2. #2
    Qwen3.8 Max

    Alibaba

    Places second on OCR-oriented visual prompts.

    Arena score1317±9
  3. #3
    Claude Opus 4.7

    Anthropic

    Completes the current leading OCR preference cluster.

    Arena score1314±7

Task · Elo

Image generation

Open ranking

For generating images from natural-language prompts.

Note: Second and third differ by one displayed Elo point.

Top 3

  1. #1
    GPT Image 2

    OpenAI · High

    Leads the selected live text-to-image preference leaderboard.

    Elo1369
  2. #2
    Reve 2.1

    Reve

    Places second on the selected live image arena.

    Elo1321
  3. #3
    Nano Banana 2 (Gemini 3.1 Flash Image Preview)

    Google

    Scores one Elo point behind second on the selected live board.

    Elo1320

Task · Arena score

Image editing

Open ranking

For editing existing images from text instructions.

Note: The two major public image-edit arenas currently disagree materially.

Top 3

  1. #1
    GPT Image 2

    OpenAI · Medium

    Leads the image-editing preference board.

    Arena score1463±4
  2. #2
    Grok Imagine Image 2.0

    xAI · Low

    Places second on the selected fresh image-edit arena.

    Arena score1439±8
  3. #3
    MAI-Image-2.6 Preview

    Microsoft AI

    Ranks third on the same image-editing board.

    Arena score1420±8

Task group · 3 tasks

Video

Open this group on LLL.bz

Task · Elo

Text → video

Open ranking

For generating complete videos from text prompts.

Note: This ranking uses the with-audio table; the no-audio table has a different order.

Top 3

  1. #1
    Wan 3.0

    Alibaba

    Leads the selected live with-audio text-to-video table.

    Elo1244
  2. #2
    Gemini Omni Flash

    Google

    Places second when native audio is included in evaluation.

    Elo1238
  3. #3
    MiniMax H3

    MiniMax

    Completes the selected with-audio text-to-video top three.

    Elo1228

Task · Elo

Image → video

Open ranking

For animating a still image into a generated video.

Note: The selected ranking uses the no-audio table to isolate visual animation quality.

Top 3

  1. #1
    Gemini Omni Flash

    Google

    Leads the selected no-audio image-to-video evaluation.

    Elo1366
  2. #2
    MiniMax H3

    MiniMax

    Places second on the selected visual-only image animation table.

    Elo1346
  3. #3
    Dreamina Seedance 2.0

    ByteDance · 720p

    Ranks third on the selected no-audio image-to-video table.

    Elo1337

Task · Elo

Video editing

Open ranking

For instruction-driven edits and transformations of existing video.

Note: The two public video-edit arenas currently disagree on first place.

Top 3

  1. #1
    Wan 3.0

    Alibaba

    Leads the selected live Artificial Analysis video-edit arena.

    Elo1192
  2. #2
    MiniMax H3

    MiniMax

    Places second on the selected live video-edit table.

    Elo1130
  3. #3
    Gemini Omni Flash

    Google

    Completes the selected live video-editing top three.

    Elo1125

Task group · 5 tasks

Audio

Open this group on LLL.bz

Task · AA-WER

Speech to text

Open ranking

For transcribing spoken audio into text.

Metric: Lower WER is better.

Top 3

  1. #1
    Fun-Realtime-ASR-preview

    Alibaba

    Has the lowest displayed AA-WER on the selected non-streaming table.

    AA-WER1.7%
  2. #2
    ElevenLabs Scribe v2

    ElevenLabs

    Places second on the selected composite transcription error metric.

    AA-WER2.2%
  3. #3
    MAI-Transcribe-1.5

    Microsoft AI

    Places third on the selected non-streaming speech benchmark.

    AA-WER2.4%

Task · Elo

Text to speech

Open ranking

For generating natural spoken audio from text.

Note: Second and third have the same displayed Elo score.

Top 3

  1. #1
    Cartesia Sonic 3.6

    Cartesia

    Leads the selected provider-voice preference arena.

    Elo1285
  2. #2
    Qwen-Audio-3.0-TTS-Plus

    Alibaba

    Shares the second displayed Elo tier on provider voices.

    Elo1240
  3. #3
    SpeechifyAI Simba 3.2

    Speechify

    Shares the same displayed Elo as second place.

    Elo1240

Task · Elo

Voice agents

Open ranking

For real-time spoken conversation with an interactive AI agent.

Note: Duplicate effort settings of Gemini 3.1 Flash Live were collapsed to the best configuration.

Top 3

  1. #1
    Gemini 3.1 Flash Live Preview

    Google · Minimal

    Has the highest distinct-model preference score in the live speech-agent arena.

    Elo1046
  2. #2
    GPT-Realtime-1.5

    OpenAI

    Is the next distinct current system after collapsing duplicate Gemini modes.

    Elo1000
  3. #3
    ElevenLabs Agents

    ElevenLabs

    Is the next distinct current voice-agent system on the selected arena.

    Elo937

Task · Elo

Music generation

Open ranking

For generating complete instrumental or vocal music.

Note: Task-level models mirror the instrumental variant. Only one specialized evaluator covers this area, so treat the order as indicative.

Instrumental

  1. #1
    Suno V5.5

    Suno

    Leads the selected instrumental preference arena.

    Elo1186
  2. #2
    Mureka V9

    Mureka

    Places second on instrumental music preference.

    Elo1174
  3. #3
    Mureka V8

    Mureka

    Ranks third on instrumental music preference.

    Elo1166

Vocals

  1. #1
    Suno V5.5

    Suno

    Leads the selected vocal-music preference arena.

    Elo1169
  2. #2
    Mureka V9

    Mureka

    Places second on the selected vocal-music arena.

    Elo1154
  3. #3
    Mureka V8

    Mureka

    Ranks third on the selected vocal-music arena.

    Elo1138

Task · Rank

Voice cloning

Open ranking

For reproducing a speaker's voice from a short reference sample.

Metric: Lower rank is better.

Note: The order comes from a single controlled-voice comparison, so only the ranking is stored, not exact scores.

Top 3

  1. #1
    Cartesia Sonic 3.6

    Cartesia

    Ranks first in the reported Controlled Voice arena update.

    Rank1
  2. #2
    Cartesia Sonic 3.5

    Cartesia

    Ranks second in the reported Controlled Voice update.

    Rank2
  3. #3
    Eleven v3

    ElevenLabs

    Ranks third in the reported Controlled Voice update.

    Rank3

Reverse lookup

All 40 models

Sorted by task wins, then total top-three appearances. This is an index, not an overall leaderboard.

Claude Fable 5

Anthropic

Wins
6
Top-3 appearances
11

General chat #1 · Russian language #1 · Expert analysis #1 · Mathematics #2 · Instruction following #2 · Copywriting #1 · Search and research #3 · PDF and documents #3 · AI agents #2 · Image understanding #1 · Text from image (OCR) #1

Claude Opus 5

Anthropic

Wins
5
Top-3 appearances
7

Expert analysis #3 · Mathematics #1 · PDF and documents #1 · Programming #1 · Web development #1 · AI agents #1 · Data analysis and SQL #2

Cartesia Sonic 3.6

Cartesia

Wins
2
Top-3 appearances
2

Text to speech #1 · Voice cloning #1

GPT Image 2

OpenAI

Wins
2
Top-3 appearances
2

Image generation #1 · Image editing #1

Suno V5.5

Suno

Wins
2
Top-3 appearances
2

Music generation #1 · Music generation #1

Wan 3.0

Alibaba

Wins
2
Top-3 appearances
2

Text → video #1 · Video editing #1

Claude Opus 4.6

Anthropic

Wins
1
Top-3 appearances
6

General chat #2 · Russian language #3 · Expert analysis #2 · Instruction following #1 · Copywriting #2 · PDF and documents #2

Qwen3.8 Max

Alibaba

Wins
1
Top-3 appearances
4

Web development #3 · Image understanding #2 · Text from image (OCR) #2 · Data analysis and SQL #1

Gemini Omni Flash

Google

Wins
1
Top-3 appearances
3

Text → video #2 · Image → video #1 · Video editing #3

Claude Opus 4.6 Search

Anthropic

Wins
1
Top-3 appearances
1

Search and research #1

Fun-Realtime-ASR-preview

Alibaba

Wins
1
Top-3 appearances
1

Speech to text #1

Gemini

Google

Wins
1
Top-3 appearances
1

Translation #1

Gemini 3.1 Flash Live Preview

Google

Wins
1
Top-3 appearances
1

Voice agents #1

Qwen3.7 Max

Alibaba

Wins
1
Top-3 appearances
1

Structured output and tool calling #1

Claude Opus 4.7

Anthropic

Wins
0
Top-3 appearances
4

General chat #3 · Instruction following #3 · Image understanding #3 · Text from image (OCR) #3

Gemini 3.7 Flash

Google

Wins
0
Top-3 appearances
3

Russian language #2 · Mathematics #3 · Copywriting #3

Kimi K3 Max

Moonshot AI

Wins
0
Top-3 appearances
3

Web development #2 · AI agents #3 · Data analysis and SQL #3

MiniMax H3

MiniMax

Wins
0
Top-3 appearances
3

Text → video #3 · Image → video #2 · Video editing #2

Mureka V8

Mureka

Wins
0
Top-3 appearances
2

Music generation #3 · Music generation #3

Mureka V9

Mureka

Wins
0
Top-3 appearances
2

Music generation #2 · Music generation #2

BTL-4

Bad Theory Labs

Wins
0
Top-3 appearances
1

Structured output and tool calling #2

Cartesia Sonic 3.5

Cartesia

Wins
0
Top-3 appearances
1

Voice cloning #2

Claude

Anthropic

Wins
0
Top-3 appearances
1

Translation #2

Dreamina Seedance 2.0

ByteDance

Wins
0
Top-3 appearances
1

Image → video #3

Eleven v3

ElevenLabs

Wins
0
Top-3 appearances
1

Voice cloning #3

ElevenLabs Agents

ElevenLabs

Wins
0
Top-3 appearances
1

Voice agents #3

ElevenLabs Scribe v2

ElevenLabs

Wins
0
Top-3 appearances
1

Speech to text #2

GPT

OpenAI

Wins
0
Top-3 appearances
1

Translation #3

GPT-5.5 Search

OpenAI

Wins
0
Top-3 appearances
1

Search and research #2

GPT-5.6 Sol

OpenAI

Wins
0
Top-3 appearances
1

Programming #2

GPT-Realtime-1.5

OpenAI

Wins
0
Top-3 appearances
1

Voice agents #2

Grok 4.6

xAI

Wins
0
Top-3 appearances
1

Programming #3

Grok Imagine Image 2.0

xAI

Wins
0
Top-3 appearances
1

Image editing #2

Ling 3.0 Flash

InclusionAI

Wins
0
Top-3 appearances
1

Structured output and tool calling #3

MAI-Image-2.6 Preview

Microsoft AI

Wins
0
Top-3 appearances
1

Image editing #3

MAI-Transcribe-1.5

Microsoft AI

Wins
0
Top-3 appearances
1

Speech to text #3

Nano Banana 2 (Gemini 3.1 Flash Image Preview)

Google

Wins
0
Top-3 appearances
1

Image generation #3

Qwen-Audio-3.0-TTS-Plus

Alibaba

Wins
0
Top-3 appearances
1

Text to speech #2

Reve 2.1

Reve

Wins
0
Top-3 appearances
1

Image generation #2

SpeechifyAI Simba 3.2

Speechify

Wins
0
Top-3 appearances
1

Text to speech #3

Interpretation contract

Methodology

Every task is ranked separately.

Public leaderboards, independent tests and official model data are weighed together.

Two modes of the same base model cannot take two places in one top 3.

Every ranking list contains exactly three distinct base models.

The complete ranking is re-checked every 10 days.

Task-level source names, URLs and measurement dates are intentionally omitted because each ranking synthesizes multiple sources.