LMSYS Chatbot Arena Leaderboard logo

LMSYS Chatbot Arena Leaderboard

New

LMSYS Chatbot Arena is a crowdsourced open plat...

Quick Facts

Pricing
Free
3.3k
views
1
favorites
Popularity Rank
#2 of 629 · by views
Added
Nov 2025
Official URL
lmarena.ai

LMArena Text Leaderboard

Source: lmarena.ai

Top 10 LLMs by current Elo / Bradley–Terry scores from LMArena human-preference battles. Click a model name to see its ToolHub family page; ↗ links to the official source.

#ModelElo Score
1Anthropic1509
2Anthropic1504
3Anthropic1502
4Anthropic1499
5Anthropic1494
6Meta1487
7Google1486
8Google1486
9Anthropic1484
10OpenAI1481
Note: this is a manually synced snapshot. For live data visit official LMArena.
ToolHub's Take
  • Elo gap from #1 to #10 is just 28 points.Differences under ~10 points are often within noise — treat the entire top 10 as one tier and pick by task fit and cost, not by who’s “smartest” on this single chart.
  • Pick by task, not by overall rank:coding, refactoring, and long-document analysis → Claude; general utility + voice + image → ChatGPT; Workspace integration + multimodal → Gemini; self-hosted or zero-cost → Meta AI / Llama.
  • Are “thinking” variants worth it? Thinking models excel at hard reasoning but cost more latency and tokens. Use the regular variant by default; switch to thinking only for tough coding, math, or multi-step problems.
  • This is the Text leaderboard only. Coding battles, web dev, vision understanding, and long context each have their own sub-arena — see WebDev / Vision / Coding tabs on the official LMArena site.

How to Read the Chatbot Arena Leaderboard

LMArena (formerly LMSYS) Chatbot Arena is the de-facto gold standard for human-preference LLM evaluation. Real users vote on blind side-by-side answers; the platform applies the Bradley–Terry model and Elo-style ratings to produce the rankings you see in the snapshot above.

The Text leaderboard captures general-chat quality. Companion leaderboards (WebDev, Vision, Coding) track domain-specific strength; if you need a model for a specific job, check the relevant sub-arena on the official site rather than defaulting to the overall top.

A few points worth knowing: Elo gaps under ~10 points are not always meaningful, "thinking" variants generally score higher but cost more latency and tokens, and newly-added models can swing rapidly before vote counts stabilize.

Related Leaderboard Resources to Compare

Official LMSYS Arena

Best when you want the source leaderboard directly and need the most current rankings without any intermediary summary.

OpenRouter Rankings

Useful if you want a more product-facing view of model availability, pricing, and ecosystem adoption alongside rankings.

Hugging Face Open LLM Leaderboards

Helpful when you want benchmark-heavy comparisons rather than crowd preference and chat-style pairwise voting.

Screenshots

LMSYS Chatbot Arena Leaderboard screenshot 1

Features

  • Crowdsourced human model comparisons
  • Elo-style LLM performance ratings
  • Bradley–Terry statistical ranking
  • Side-by-side anonymous chat battles
  • Unified view of open and closed models
  • Continuously updated leaderboard data
  • Free web-based evaluation access
  • Granular insights across task types

Tags

AI
artificial-intelligence
lmsys
chatbot

How to Use the Chatbot Arena Effectively

  • Check the "Coding" or "Hard Prompts" category leaderboards specifically if you are looking for a model to handle complex logic or software development.

  • Participate in "Side-by-side" battles to contribute to the ELO rankings while testing your specific edge-case prompts against two anonymous models.

  • Monitor the "Style Control" and "Long Context" updates to see which models excel at following strict formatting or handling massive documents.

Frequently Asked Questions

User Reviews

No reviews yet. Be the first to share your experience!