The model that wins your task, wired into your business.

No single model wins everything. Today the leader in code is not the leader in voice, the best model for documents loses on images, and half of these positions move within a month. That is what the board below is for: the top 3 for each of 26 tasks, rebuilt every 10 days from public leaderboards and independent tests.

Then I build the thing that uses it. Agents that handle support, pipelines that read invoices and contracts, voice that answers the phone, bots that qualify leads while you sleep. It plugs into what you already run — CRM, spreadsheets, Telegram, your own API — and every step calls the model that wins that particular step, whichever provider it belongs to.

Write me on Telegram — @simvimTell me which process eats your week. I answer myself.

Text

Code

Images

Video

Audio

General chat

For broad conversational quality across everyday prompts.

The primary ranking is a blind-preference arena; confidence intervals overlap.

#1

Claude Fable 5

Anthropic

Ranks first in the current broad text preference arena.

Arena score1508±5

  • broad chat preference
  • strong user ratings
#2

Claude Opus 4.6

Anthropic

A near-tied option for broad conversational prompts.

Arena score1504±4

  • broad conversation quality
  • stable arena score
#3

Claude Opus 4.7

Anthropic

Places inside the same tightly clustered chat tier.

Arena score1502±4

  • broad prompt handling
  • competitive preference score

Russian language

For Russian-language conversation and writing.

Fresh primary data exists, but three independent Russian-specific sources reproducing the full top three were not found.

#1

Claude Fable 5

Anthropic

Leads the current Russian-language Arena slice.

Arena score1518±13

  • russian chat preference
  • current arena lead
#2

Gemini 3.7 Flash

Google

Preliminary

Scores effectively level with the current leader in Russian.

Arena score1517±24

  • strong russian preference
  • near leader score
#3

Claude Opus 4.6

Anthropic

Remains inside the leading Russian preference cluster.

Arena score1509±8

  • russian prompt quality
  • stable arena placement

Expert analysis

For difficult analytical and professional knowledge tasks.

The top three intervals overlap, so the exact order is not statistically decisive.

#1

Claude Fable 5

Anthropic

Has the highest current score on expert-oriented Arena prompts.

Arena score1548±12

  • expert prompt preference
  • strong reasoning results
#2

Claude Opus 4.6

Anthropic

Scores essentially level with first place on expert prompts.

Arena score1546±9

  • complex prompt handling
  • high expert preference
#3

Claude Opus 5

Anthropic

Places in the leading cluster for specialist analysis.

Arena score1541±12

  • complex analytical prompts
  • strong knowledge work

Mathematics

For hard calculations, proofs and multi-step maths.

All three reported confidence intervals overlap materially.

#1

Claude Opus 5

Anthropic

Has the highest displayed score in the current Math Arena slice.

Arena score1545±24

  • hard math prompts
  • multi-step reasoning
#2

Claude Fable 5

Anthropic

Ranks inside the same leading hard-math cluster.

Arena score1528±17

  • strong math reasoning
  • competitive proof solving
#3

Gemini 3.7 Flash

Google

Preliminary

Posts a near-tied preliminary score on difficult mathematics.

Arena score1523±33

  • competitive math score
  • multi-step problem solving

Instruction following

For prompts with detailed constraints and required formats.

First and second are practically tied within the reported intervals.

#1

Claude Opus 4.6

Anthropic

Edges the current instruction-following Arena by displayed score.

Arena score1514±5

  • constraint following quality
  • structured prompt handling
#2

Claude Fable 5

Anthropic

Scores effectively level with first place on constrained prompts.

Arena score1512±7

  • complex constraint handling
  • strong prompt adherence
#3

Claude Opus 4.7

Anthropic

Maintains a high preference score on detailed instructions.

Arena score1503±6

  • detailed prompt handling
  • format constraint adherence

Copywriting

For persuasive, creative and polished marketing text.

The source measures blind writing preference, not measured conversion performance.

#1

Claude Fable 5

Anthropic

Ranks first on current blind creative-writing preference.

Arena score1508±9

  • creative writing preference
  • natural prose quality
#2

Claude Opus 4.6

Anthropic

Scores close to first place for creative prose.

Arena score1499±7

  • polished longform prose
  • strong writing preference
#3

Gemini 3.7 Flash

Google

Preliminary

Places in the current leading creative-writing cluster.

Arena score1493±18

  • creative prompt handling
  • competitive writing preference

Search and research

For web search, evidence gathering and sourced answers.

The dedicated search board updates slower than the ten-day rebuild cycle.

#1

Claude Opus 4.6 Search

Anthropic

Leads the dedicated search-and-retrieval board.

Arena score1253±5

  • search answer preference
  • web retrieval workflow
#2

GPT-5.5 Search

OpenAI

Places second on the same search board.

Arena score1240±5

  • search focused responses
  • strong arena preference
#3

Claude Fable 5

Anthropic

Completes the leading trio for search tasks.

Arena score1237±8

  • search answer quality
  • broad research reasoning

PDF and documents

For understanding long documents, PDFs and mixed document content.

The document board updates slower than the ten-day cycle; duplicate Claude Opus 4.6 configurations are collapsed into one entry.

#1

Claude Opus 5

Anthropic

Leads the dedicated document-handling board.

Arena score1520±15

  • document prompt preference
  • long context reasoning
#2

Claude Opus 4.6

Anthropic

Ranks second after collapsing duplicate configurations.

Arena score1510±6

  • document understanding quality
  • stable arena score
#3

Claude Fable 5

Anthropic

Becomes third after enforcing one configuration per base model.

Arena score1504±9

  • document comprehension quality
  • long prompt handling

Translation

For translating text accurately between languages.

Measured by model family rather than exact version: the evaluation reports Gemini, Claude and GPT as families, so read this as a family-level ranking.

#1

Gemini

Google

Ranks first for translation quality across the tested sample.

AQI77.7

  • multilingual translation quality
  • broad language coverage
#2

Claude

Anthropic

Ranks second across the same translation sample.

AQI75.6

  • multilingual translation quality
  • broad project coverage
#3

GPT

OpenAI

Ranks third across the same translation sample.

AQI73.1

  • multilingual translation coverage
  • broad language testing
How the ranking is built

There is no single “best AI model” here. Every task is judged on its own: public leaderboards, independent tests and official model data, weighed together rather than copied from one board. The whole thing is re-checked every 10 days. Two modes of the same base model never take two places in one top 3. Likes are a separate reader signal — they never move the ranking.