LMArena (Chatbot Arena)

From Systems Analysis Wiki
Jump to navigation Jump to search

Arena AI (until January 28, 2026, LMArena (Large Model Arena); formerly Chatbot Arena) is an open, web-based platform for the crowdsourced evaluation and comparison of large language models (LLMs) and multimodal models (text, image, video) based on real-world human preferences. At its core are anonymous pairwise comparisons of models (blind battles) and the Elo rating system; from these, the platform publishes public, transparent leaderboards that are considered one of the most recognized independent benchmarks for frontier AI models.[1][2]

The platform emerged in 2023 as an academic research project of LMSYS Org (Large Model Systems Organization) at the University of California, Berkeley (Sky Computing Lab). In September 2024, the project "graduated" to its own domain, lmarena.ai ("Graduation")[3]. In April 2025, the independent company Arena Intelligence Inc. was incorporated, and on May 21, 2025, it raised a $100 million seed round (a $600 million valuation; lead investors a16z and UC Investments)[4][5]. On January 6, 2026, the company closed a $150 million Series A round at a $1.7 billion post-money valuation, led by Felicis and UC Investments[6][7]. On January 28, 2026, the platform underwent its final rebrand: it was renamed Arena and moved to the domain arena.ai[8][9].

History

The platform launched in April 2023 under the name Chatbot Arena, as a research project of LMSYS Org (Large Model Systems Organization) at the University of California, Berkeley (Sky Computing Lab). It was one of the first tools for the crowdsourced evaluation of large language models through anonymous pairwise comparisons (blind battles) and the Elo rating system, based on real user preferences.

  • April 24, 2023 — technical launch of Chatbot Arena.
  • May 3, 2023 — official public launch and release of the first leaderboard.
  • 2023 — release of the first open datasets: 33K paired dialogues (July) and LMSYS-Chat-1M (September, around 1 million real dialogues).
  • March 1, 2024 — publication of the platform's official policy and formalization of its mission as an open, community-driven evaluation system.
  • June 27, 2024 — addition of image support and the start of expansion into multimodal tasks.
  • September 20, 2024 — "Graduation": move to the independent domain lmarena.ai.
  • Late 2024 – spring 2025 — launch of specialized arenas (Arena-Hard, WebDev Arena, RepoChat Arena, Style/Sentiment Control, etc.).
  • April 17, 2025 — official incorporation as the independent company Arena Intelligence Inc. and launch of the beta of the revamped platform under the LMArena brand.
  • May 21, 2025 — announcement of the company's formation and a $100 million seed round ($600 million valuation).
  • July 31, 2025 — release of an open dataset of 140K recent Text Arena dialogues.
  • December 18, 2025 — release of Arena-Rank, the open-source package implementing the ranking methodology.
  • January 6, 2026 — close of a $150 million Series A round at a $1.7 billion post-money valuation.
  • January 2026 — launch of Video Arena (full support for the video modality).
  • January 28, 2026 — final rebrand: the platform was renamed Arena and moved to the domain arena.ai.

As of March 2026, Arena (arena.ai) serves more than 5 million monthly active users across more than 150 countries, has accumulated tens of millions of votes, and remains one of the most recognized independent tools for evaluating frontier AI models based on real-world human preferences.

How the evaluation works

The user enters a query (prompt) and receives two responses from randomly selected anonymous models ("A" and "B"), then votes for the better response (or declares a tie or that neither is satisfactory). Ranking is based on the Bradley-Terry statistical model (a logistic regression over pairwise preferences), conceptually close to Elo[1]. The platform publishes the Arena Score and confidence intervals, and applies sampling corrections (re-weighting) to remain unbiased under non-uniform sampling[10].

Transparency and openness. The evaluation and ranking pipelines are open source: the original infrastructure lives in the FastChat repository[11], and in December 2025 the leaderboard methodology was released as a standalone Python package, Arena-Rank —which now powers all of the site's leaderboards and is roughly 30 times faster than the FastChat-based version[12]. The platform also periodically releases portions of the raw data for verification and research (for example, the release of 140K conversations in July 2025)[10][13]. Per the FAQ and the warnings on the homepage, user queries may be shared with model providers and partially published for research purposes, so sensitive data should not be submitted[14][15].

Selection and sampling rules. The leaderboards include publicly available models (open weights, public API, or public service). Stabilizing a score typically requires ≥1,000 votes; at least 20% of battles are fought solely between public models; sampling probability increases with rating and uncertainty, and the re-weighted regression ensures that the final scores remain unbiased[10].

Automatic metrics and style control. To speed up evaluation and reduce the effects of "style" preferences, auxiliary methodologies are used: MT-Bench (LLM-as-a-judge)[16], Arena-Hard (automatic generation of difficult questions)[17], and Style/Sentiment Control (modeling and correcting the effect of tone/sentiment on preferences)[18]. For Arena-Hard-Auto, a very high agreement with live human votes has been reported (up to about 98.6% under controlled conditions)[19].

Arenas and evaluation domains

The platform has evolved into a set of "arenas" by task type:

  • Text Arena — general conversations and tasks; the main leaderboard[20].
  • Vision Arena — multimodal "text→image/video/image analysis" models[21].
  • Text-to-Image and Image Edit — image generation and editing (including the nano-banana case)[22][23].
  • Text-/Image-to-Video — video generation[24].
  • WebDev Arena — building web applications from descriptions[25].
  • RepoChat Arena — AI engineering tasks involving code and repositories[26].
  • Search Arena — models with web-search connectivity; first launched in April 2025 (legacy), then migrated to the main site, accompanied by a dataset and a publication[27][28][29].
  • BiomedArena.AI — domain-specific evaluation for biomedical tasks (in partnership with DataTecnica)[30].

Application and impact

  • Industry showcase. Major providers (OpenAI, Anthropic, Google, etc.) regularly test and showcase their models on the platform; industry media describe it as an important benchmark[5][31]. In a NAACL-2025 industry paper, the Chatbot Arena Elo score is described as the "gold industry-standard"[32].
  • Pre-release testing. The policy allows anonymous previews of "unreleased" models, with notice to the community and subsequent publication of public evaluations after release; a minimum of ≈1,000 votes is required for stabilization[10].
  • Notable episodes. In spring 2025, the anonymous model Llama-4 Maverick-03-26-Experimental was discussed (an incident involving its comparison with the public versions), which drew wide press attention and prompted updates to the rules and communications[33][34]. In August 2025, "nano-banana" was revealed to be Gemini 2.5 Flash Image and took the top positions in the visual arenas[23][22].

Limitations and criticism

Despite its scale and popularity, the approach has limitations:

  • Subjectivity and style effects. Voting preferences depend on the tone and form of the response; the team is implementing Style/Sentiment Control to decouple "style" from "content"[18].
  • Lack of audience representativeness. The active core consists of tech enthusiasts and developers; for domain-specific scenarios, specialized arenas are created (Search, WebDev, Biomed, etc.)[35].
  • Vulnerability to manipulation and bias. Research from 2025 shows that, without strict defenses, vote-rigging strategies involving hundreds to thousands of votes are possible; however, collaboration between researchers and LMArena led to protective measures (CAPTCHA, login, bot protection, anomaly detection) and an increased "cost of attack"[36][37][38].
  • Methodological criticism. The paper The Leaderboard Illusion (April 2025) points to systematic and institutional factors that can distort the competitive landscape; LMArena published a detailed response and maintains a public changelog of its methodology[39][40][41].

Bibliography

  • Chiang, W.-L. et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv:2403.04132.
  • Zheng, L. et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685.
  • Li, T. et al. (2024). From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline. arXiv:2406.11939.
  • Ameli, S.; Zhuang, S.; Stoica, I.; Mahoney, M. W. (2024). A Statistical Framework for Ranking LLM-Based Chatbots. arXiv:2412.18407.
  • Boubdir, M. et al. (2023). Elo Uncovered: Robustness and Best Practices in Language Model Evaluation. arXiv:2311.17295.
  • Huang, J. Y.; Shen, Y.; Wei, D.; Broderick, T. (2025). Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings. arXiv:2508.11847.
  • Xu, Y.; Ruis, L.; Rocktäschel, T.; Kirk, R. (2025). Investigating Non-Transitivity in LLM-as-a-Judge. arXiv:2502.14074.
  • Li, H. et al. (2024). LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods. arXiv:2412.05579.
  • Zheng, L. et al. (2024). LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset. arXiv:2309.11998.
  • Dubois, Y. et al. (2024). Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. arXiv:2404.04475.
  • Singh, S. et al. (2025). The Leaderboard Illusion. arXiv:2504.20879.
  • Min, R.; Pang, T.; Du, C.; Liu, Q.; Cheng, M.; Lin, M. (2025). Improving Your Model Ranking on Chatbot Arena by Vote Rigging. arXiv:2501.17858.

References

  1. 1.0 1.1 Chiang, W.-L. et al. "Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference." arXiv:2403.04132, 2024. arXiv
  2. "Hello from LMArena: The Community Platform for Exploring Frontier AI." LMArena Blog, June 23, 2025. [1]
  3. "Announcing a New Site for Chatbot Arena." LMSYS Blog, September 20, 2024. [2]
  4. "LMArena Secures $100M in Seed Funding to Bring Scientific Rigor to AI Reliability." PR Newswire, May 21, 2025. [3]
  5. 5.0 5.1 Wiggers, K. "LM Arena, the organization behind popular AI leaderboards, lands $100M." TechCrunch, May 21, 2025. [4]
  6. "LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform." PR Newswire, January 6, 2026. [5]
  7. "LMArena lands $1.7B valuation four months after launching its product." TechCrunch, January 6, 2026. [6]
  8. "LMArena is now Arena." Arena Blog, January 28, 2026. [7]
  9. "LMArena is Growing to Support our Community Platform." LMArena Blog, April 17, 2025.
  10. 10.0 10.1 10.2 10.3 Arena Leaderboard Policy. Arena Blog, last updated April 30, 2026. [8]
  11. lm-sys/FastChat (GitHub). [9]
  12. "Arena-Rank: Open Sourcing the Leaderboard Methodology." Arena Blog, December 18, 2025. [10]. Repository: lmarena/arena-rank.
  13. Y. Song. "A Deep Dive into Recent Arena Data." LMArena Blog, July 31, 2025. [11]
  14. FAQ. Arena. [12]
  15. Arena homepage (disclaimer about possible data publication and transfer to providers). [13]
  16. Zheng, L. et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." arXiv:2306.05685, 2023. [14]
  17. Li, T. et al. "From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline." arXiv:2406.11939, 2024. [15]
  18. 18.0 18.1 "Does Sentiment Matter Too? Introducing Sentiment Control." LMArena Blog, April 22, 2025. [16]
  19. Li, T. et al. "From Crowdsourced Data…" arXiv:2406.11939 (agreement tables). [17]
  20. Text Arena (English). Arena. [18]
  21. Vision Arena. Arena. [19]
  22. 22.0 22.1 Text-to-Image Arena. Arena. [20]
  23. 23.0 23.1 "Nano-Banana (Gemini 2.5 Flash Image): Try it on LMArena." LMArena Blog, August 27, 2025. [21]
  24. Text-to-Video and Image-to-Video Leaderboards. Arena. [22] [23]
  25. "WebDev Arena: A Live LLM Leaderboard for Web App Development." LMArena Blog, March 10, 2025. [24]
  26. "RepoChat Arena: A Live Benchmark for AI Software Engineers." LMArena Blog, April 9, 2025. [25]
  27. "Introducing the Search Arena." LMArena Blog, April 14, 2025. [26]
  28. "Search Arena & What We're Learning About Human Preference." LMArena Blog, July 23, 2025. [27]
  29. Frick, E. et al. "Search Arena: Analyzing Search-Augmented LLMs." arXiv:2506.05334, 2025. [28]
  30. "Introducing BiomedArena.AI." LMArena Blog, August 19, 2025. [29]
  31. Google. "Gemma 3…," March 12, 2025 (link to LMArena results). [30]
  32. Spangher, L. et al. "Chatbot Arena Estimate…." NAACL Industry, 2025. [31]
  33. "Meta's experimental Llama 4 model briefly topped AI leaderboard…." The Register, April 7, 2025. [32]
  34. LMArena's official clarifications and posts on X about the incident (April 2025). [33]
  35. "Search Arena & What We're Learning…." LMArena Blog, July 23, 2025. [34]
  36. Min, R. et al. "Improving Your Model Ranking on Chatbot Arena by Vote Rigging." arXiv:2501.17858, 2025. [35]
  37. Huang, Y. et al. "Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards." arXiv:2501.07493, 2025. [36]
  38. "Hundreds of rigged votes can skew…." Fast Company, February 6, 2025. [37]
  39. Singh, S. et al. "The Leaderboard Illusion." arXiv:2504.20879, 2025. [38]
  40. "Our Response to 'The Leaderboard Illusion'." LMArena Blog, May 9, 2025. [39]
  41. Leaderboard Changelog. Arena Blog. [40]