Qwen3.8 Max
Alibaba
Leads the fresh Data & Analytics WebDev slice.
Arena score1614±30
- data app generation
- analytics interface quality
No single model wins everything. Today the leader in code is not the leader in voice, the best model for documents loses on images, and half of these positions move within a month. That is what the board below is for: the top 3 for each of 26 tasks, rebuilt every 10 days from public leaderboards and independent tests.
Then I build the thing that uses it. Agents that handle support, pipelines that read invoices and contracts, voice that answers the phone, bots that qualify leads while you sleep. It plugs into what you already run — CRM, spreadsheets, Telegram, your own API — and every step calls the model that wins that particular step, whichever provider it belongs to.
For tables, CSV analysis, SQL generation and data-driven conclusions.
The fresh source measures data-oriented web applications, not normalized text-to-SQL execution accuracy. BIRD and Spider2 did not yield a clean current raw-model top three that could be reconciled with this board.
Alibaba
Leads the fresh Data & Analytics WebDev slice.
Arena score1614±30
Anthropic
Places effectively level with first on data-oriented web tasks.
Arena score1612±16
Moonshot AI
Completes the leading Data & Analytics WebDev cluster.
Arena score1600±22
There is no single “best AI model” here. Every task is judged on its own: public leaderboards, independent tests and official model data, weighed together rather than copied from one board. The whole thing is re-checked every 10 days. Two modes of the same base model never take two places in one top 3. Likes are a separate reader signal — they never move the ranking.