The model that wins your task, wired into your business.

No single model wins everything. Today the leader in code is not the leader in voice, the best model for documents loses on images, and half of these positions move within a month. That is what the board below is for: the top 3 for each of 26 tasks, rebuilt every 10 days from public leaderboards and independent tests.

Then I build the thing that uses it. Agents that handle support, pipelines that read invoices and contracts, voice that answers the phone, bots that qualify leads while you sleep. It plugs into what you already run — CRM, spreadsheets, Telegram, your own API — and every step calls the model that wins that particular step, whichever provider it belongs to.

Write me on Telegram — @simvimTell me which process eats your week. I answer myself.

Text

Code

Images

Video

Audio

Programming

For repository work, terminal coding and software engineering.

Terminal-Bench measures agentic terminal coding; SWE-bench and Code Arena measure different coding workloads.

#1

Claude Opus 5

Anthropic

Has the highest reported Terminal-Bench result of the three.

Terminal-Bench score89.1%

  • agentic terminal coding
  • repository task solving
#2

GPT-5.6 Sol

OpenAI

Trails the reported leader by three tenths of a point.

Terminal-Bench score88.8%

  • agentic coding tasks
  • terminal workflow quality
#3

Grok 4.6

xAI

Places inside the same tightly grouped Terminal-Bench frontier.

Terminal-Bench score88.4%

  • terminal task solving
  • competitive coding score

Web development

For building and editing interactive web interfaces.

The ranking measures blind preference over generated web applications.

#1

Claude Opus 5

Anthropic

Leads the current WebDev Arena by displayed score.

Arena score1691±9

  • web interface generation
  • strong user preference
#2

Kimi K3 Max

Moonshot AI

Places second on blind evaluation of generated web applications.

Arena score1674±11

  • frontend generation quality
  • competitive arena score
#3

Qwen3.8 Max

Alibaba

Preliminary

Places close behind second in the current WebDev Arena.

Arena score1669±13

  • web app generation
  • competitive frontend preference

AI agents

For autonomous multi-step work with tools and external environments.

Higher net improvement is better.

Claude Opus 5 Max is omitted because it is another configuration of the same base model.

#1

Claude Opus 5

Anthropic

Has the highest distinct-base-model result in the current Agent Arena.

Net improvement12.47%±1.54

  • multi-step agent work
  • tool workflow quality
#2

Claude Fable 5

Anthropic

Ranks second after collapsing duplicate Opus 5 configurations.

Net improvement11.57%±1.70

  • agent task execution
  • multi-step tool use
#3

Kimi K3 Max

Moonshot AI

Is the next distinct base model in the Agent Arena ranking.

Net improvement10.41%±0.62

  • agentic workflow quality
  • tool task execution

Data analysis and SQL

For tables, CSV analysis, SQL generation and data-driven conclusions.

The fresh source measures data-oriented web applications, not normalized text-to-SQL execution accuracy. BIRD and Spider2 did not yield a clean current raw-model top three that could be reconciled with this board.

#1

Qwen3.8 Max

Alibaba

Preliminary

Leads the fresh Data & Analytics WebDev slice.

Arena score1614±30

  • data app generation
  • analytics interface quality
#2

Claude Opus 5

Anthropic

Places effectively level with first on data-oriented web tasks.

Arena score1612±16

  • data workflow generation
  • analytical app building
#3

Kimi K3 Max

Moonshot AI

Completes the leading Data & Analytics WebDev cluster.

Arena score1600±22

  • data app generation
  • analytics workflow quality

Structured output and tool calling

For reliable function calls, JSON arguments and structured tool use.

Scores come from the Berkeley function-calling benchmark. Tool schemas differ between providers, so narrow gaps here mean little in practice.

#1

Qwen3.7 Max

Alibaba

Has the highest BFCL V4 score of the three.

BFCL V4 score75.0%

  • function call accuracy
  • structured tool output
#2

BTL-4

Bad Theory Labs

Places second on the captured BFCL V4 score view.

BFCL V4 score73.5%

  • function call reliability
  • structured argument handling
#3

Ling 3.0 Flash

InclusionAI

Places third on the captured BFCL V4 score view.

BFCL V4 score73.0%

  • function call accuracy
  • json argument handling
How the ranking is built

There is no single “best AI model” here. Every task is judged on its own: public leaderboards, independent tests and official model data, weighed together rather than copied from one board. The whole thing is re-checked every 10 days. Two modes of the same base model never take two places in one top 3. Likes are a separate reader signal — they never move the ranking.