AI Model Leaderboard

16 LLMs ranked by benchmark performance & API pricing. Updated daily.

Full Benchmark Breakdown

All scores from public leaderboards. Higher is better.

ModelMMLUHumanEvalSWE-benchGSM8KGPQAIFEvalELO
✧Gemini90%88.4%36.1%95.8%62.2%84.1%1301
✦Muse Glimmer90.5%91%42%96.8%58%86.5%1295
✦ChatGPT88.7%90.2%33.2%95.8%53.6%85.6%1287
✸Claude89.3%93.7%49%96.4%59.4%89.3%1271
✕Grok86%85%30%94%48%—1268
◈DeepSeek88.5%82.6%38.8%97.3%51.1%78.3%1257
☾Kimi K287%85%35%95%50%—1255
◑Llama 486%81%28%94%45%—1228
◈Tongyi Qianwen85%80%—93%44%—1220
◆Mistral Large81%81%24%90%38%—1215
◬Amazon Nova85%80%—92%45%—1210
▣GLM-584%78%—92%42%—1205
◈Reka Core83%78%—90%40%—1200
⬢Nemotron82%75%—88%38%—1190
ϕPhi-479%76%—88%35%—1185
▦IBM Granite78%70%—85%32%—1170

Compare models side-by-side

Use our interactive comparison deck with weighted scoring.

Compare AI Tools