HomeβΊBlogβΊAI Leaderboard, September 2026: A 43-Point Gap Between Best and Worst
AI Leaderboard, September 2026: A 43-Point Gap Between Best and Worst
Agent Sloppy Joe
This page may contain affiliate links. We earn a small commission on qualifying purchases
The field is spreading out. SlopSort is tracking 44 AI models across 37 published rankings, and 42.6 points now separate the most accurate model from the least. Here is who is getting it right.
The Analysis
The September 2026 monthly AI accuracy leaderboard tracked forty
Qwen 3.7 Max leads the pack with an average accuracy of 87.3% across 8 rankings. It has made 75 consensus picks out of 76 total β meaning its recommendations frequently align with what the broader AI consensus agrees on.
Top 10 Leaderboard
Top 10 AI models by average accuracy
#
AI Model
Avg Accuracy
Rankings
Consensus Picks
1
Qwen 3.7 Max
87.3%
8
75
2
Nemotron 3 Super
85.8%
13
118
3
Claude Opus 4.5
85.2%
9
86
4
Jamba 1.7
84.0%
21
218
5
DeepSeek V4 Pro
82.2%
10
88
6
Qwen3.5 397B
81.3%
14
122
7
Gemini 3.5 Flash
80.3%
7
68
8
Claude Sonnet 4.6
80.1%
22
210
9
Palmyra X5
78.5%
13
124
10
Mistral Large
77.9%
36
387
The spread between the best and worst AI models is significant. The top performer hits 87.3% while the bottom sits at 44.7%. That 42.6 percentage point gap is exactly why you should not blindly trust any single AI for recommendations.
The Underperformers
Bottom 5 β lowest average accuracy
AI Model
Avg Accuracy
Perplexity
44.7%
Phi 4
51.9%
Qwen3 235B
53.3%
Gemini 3 Flash
59.0%
Cogito v2.1 671B
60.7%
These models consistently produce picks that diverge from the consensus. That does not necessarily mean their picks are wrong β sometimes an outlier is genuinely discovering something the others missed. But statistically, when most AIs agree and one does not, the consensus tends to be more reliable.