The model that wins your task, wired into your business.

No single model wins everything. Today the leader in code is not the leader in voice, the best model for documents loses on images, and half of these positions move within a month. That is what the board below is for: the top 3 for each of 26 tasks, rebuilt every 10 days from public leaderboards and independent tests.

Then I build the thing that uses it. Agents that handle support, pipelines that read invoices and contracts, voice that answers the phone, bots that qualify leads while you sleep. It plugs into what you already run — CRM, spreadsheets, Telegram, your own API — and every step calls the model that wins that particular step, whichever provider it belongs to.

Write me on Telegram — @simvimTell me which process eats your week. I answer myself.

Text

Code

Images

Video

Audio

Speech to text

For transcribing spoken audio into text.

Lower WER is better.

#1

Fun-Realtime-ASR-preview

Alibaba

Preview

Has the lowest displayed AA-WER on the selected non-streaming table.

AA-WER1.7%

  • lowest displayed wer
  • strong transcription accuracy
#2

ElevenLabs Scribe v2

ElevenLabs

Places second on the selected composite transcription error metric.

AA-WER2.2%

  • low transcription error
  • broad speech recognition
#3

MAI-Transcribe-1.5

Microsoft AI

Places third on the selected non-streaming speech benchmark.

AA-WER2.4%

  • low speech error
  • strong accuracy latency

Text to speech

For generating natural spoken audio from text.

Second and third have the same displayed Elo score.

#1

Cartesia Sonic 3.6

Cartesia

Leads the selected provider-voice preference arena.

Elo1285

  • natural speech preference
  • streaming voice quality
#2

Qwen-Audio-3.0-TTS-Plus

Alibaba

Shares the second displayed Elo tier on provider voices.

Elo1240

  • natural voice generation
  • competitive speech preference
#3

SpeechifyAI Simba 3.2

Speechify

Shares the same displayed Elo as second place.

Elo1240

  • natural speech generation
  • strong provider voices

Voice agents

For real-time spoken conversation with an interactive AI agent.

Duplicate effort settings of Gemini 3.1 Flash Live were collapsed to the best configuration.

#1

Gemini 3.1 Flash Live Preview

Google

Preview

Has the highest distinct-model preference score in the live speech-agent arena.

Elo1046

  • live conversational speech
  • strong agent preference
#2

GPT-Realtime-1.5

OpenAI

Is the next distinct current system after collapsing duplicate Gemini modes.

Elo1000

  • realtime spoken dialogue
  • interactive voice workflow
#3

ElevenLabs Agents

ElevenLabs

Is the next distinct current voice-agent system on the selected arena.

Elo937

  • voice agent platform
  • realtime spoken interaction

Music generation

For generating complete instrumental or vocal music.

Task-level models mirror the instrumental variant. Only one specialized evaluator covers this area, so treat the order as indicative.

#1

Suno V5.5

Suno

Leads the selected instrumental preference arena.

Elo1186

  • instrumental music preference
  • current arena lead
#2

Mureka V9

Mureka

Places second on instrumental music preference.

Elo1174

  • instrumental generation quality
  • strong music preference
#3

Mureka V8

Mureka

Ranks third on instrumental music preference.

Elo1166

  • instrumental music quality
  • competitive arena score

Voice cloning

For reproducing a speaker's voice from a short reference sample.

Lower rank is better.

The order comes from a single controlled-voice comparison, so only the ranking is stored, not exact scores.

#1

Cartesia Sonic 3.6

Cartesia

Ranks first in the reported Controlled Voice arena update.

Rank1

  • controlled voice cloning
  • streaming cloned speech
#2

Cartesia Sonic 3.5

Cartesia

Ranks second in the reported Controlled Voice update.

Rank2

  • controlled voice cloning
  • reference voice matching
#3

Eleven v3

ElevenLabs

Ranks third in the reported Controlled Voice update.

Rank3

  • reference voice cloning
  • expressive cloned speech
How the ranking is built

There is no single “best AI model” here. Every task is judged on its own: public leaderboards, independent tests and official model data, weighed together rather than copied from one board. The whole thing is re-checked every 10 days. Two modes of the same base model never take two places in one top 3. Likes are a separate reader signal — they never move the ranking.