# LMSYS

LMSYS is an open research organization focused on large language models, known for creating the Chatbot Arena evaluation platform and the Vicuna model family.

LMSYS (Large Model Systems Organization) is an open research collaboration that develops and evaluates large language models (LLMs). Founded in 2023 by researchers from the University of California, Berkeley, Carnegie Mellon University, Stanford University, and the University of California, San Diego, LMSYS aims to advance the field of artificial intelligence through open-source tools, public benchmarks, and community-driven research. The organization is best known for its Chatbot Arena, a crowdsourced platform for comparing LLMs, and for releasing the Vicuna series of models, which have become widely used in both academia and industry.

LMSYS operates with a mission to make large model research accessible and transparent. Its projects emphasize real-world evaluation, reproducibility, and the democratization of AI technologies. By combining academic rigor with practical deployment, LMSYS has become a central hub for LLM benchmarking and development, influencing how models are assessed and improved globally.

## History and Founding

LMSYS was established in early 2023, emerging from collaborative efforts among researchers who recognized the need for a shared infrastructure to evaluate and develop large models. The founding institutions include the University of California, Berkeley, Carnegie Mellon University, Stanford University, and the University of California, San Diego. The initiative was partly inspired by the rapid proliferation of open-source LLMs following the release of models like LLaMA by Meta. The founders aimed to create a platform where models could be compared fairly and where research findings could be shared openly.

The organization quickly gained traction within the AI community, attracting contributions from volunteers and researchers worldwide. Its first major public release was the Chatbot Arena in April 2023, which allowed users to interact with anonymous models and vote on their responses. This crowdsourced approach to evaluation was novel and provided a scalable alternative to traditional static benchmarks.

## Chatbot Arena

Chatbot Arena is LMSYS's flagship project, launched in April 2023. It is a web-based platform where users can submit prompts and receive responses from two anonymous LLMs. Users then vote on which response is better, contributing to a dynamic leaderboard based on the Elo rating system, similar to chess rankings. The platform supports a wide range of models, including proprietary ones like GPT-4 and Claude, as well as open-source models like Llama and Mistral.

The Arena's strength lies in its real-world, human-preference-based evaluation, which captures aspects of model quality that automated metrics often miss. As of early 2025, the Arena has collected over two million votes and has become a de facto standard for LLM comparison. It also provides a public API and dataset, enabling researchers to analyze user preferences and model performance. The leaderboard is updated regularly, reflecting the fast-paced evolution of LLMs.

## Vicuna Models

Vicuna is a family of open-source chat models developed by LMSYS, first released in March 2023. The initial Vicuna-13B was created by fine-tuning Meta's LLaMA-13B on approximately 70,000 user-shared conversations gathered from ShareGPT. The training process used a modified version of the Alpaca training pipeline, with optimizations for memory and speed. Vicuna-13B demonstrated impressive performance, achieving over 90% of the quality of ChatGPT in human evaluations, according to LMSYS's internal assessments.

Subsequent versions include Vicuna-7B and Vicuna-33B, and later iterations based on newer base models. Vicuna models are widely used in research and production due to their strong performance-to-size ratio and permissive licenses. They have also served as baselines in numerous academic studies and have been integrated into various applications, including chatbots and virtual assistants.

## Other Projects and Contributions

Beyond Chatbot Arena and Vicuna, LMSYS has contributed to several other research initiatives. These include:

- **MT-Bench**: A multi-turn benchmark designed to evaluate chat models on more complex, conversational tasks. MT-Bench consists of 80 multi-turn questions covering topics like writing, reasoning, and coding, and is scored by GPT-4 as a judge.
- **LongBench**: A benchmark for long-context understanding, testing models on tasks that require processing documents of up to 10,000 words.
- **Chatbot Arena Dataset**: A large-scale dataset of human preferences, released to the research community for further analysis.
- **Model Evaluation Tools**: LMSYS has developed open-source tools for efficient model serving and evaluation, such as FastChat, a platform for training, serving, and evaluating chatbots.

These projects have collectively advanced the understanding of LLM capabilities and limitations, particularly in areas like alignment, robustness, and efficiency.

## Impact and Reception

LMSYS has had a significant impact on the AI ecosystem. Chatbot Arena has become a trusted reference for model performance, influencing both developers and users. The Vicuna models have been downloaded millions of times and have inspired numerous derivative works. The organization's emphasis on open research has fostered a collaborative culture, with many external contributors participating in its projects.

Critics have noted potential biases in crowdsourced evaluations, such as the influence of model verbosity or style on user preferences. LMSYS has acknowledged these issues and continues to refine its methodology, including the use of statistical techniques to mitigate biases. Despite these challenges, the organization's contributions are widely regarded as valuable for the community.

## Future Directions

Looking ahead, LMSYS plans to expand its evaluation platforms to cover multimodal models, including vision-language models. The organization is also exploring ways to make evaluation more efficient and scalable, possibly through automated judging and adaptive testing. Additionally, LMSYS aims to foster more collaborations with industry partners and other academic institutions to accelerate research progress.

As of early 2025, LMSYS remains at the forefront of LLM research, with ongoing projects and a growing community. Its open-source philosophy and commitment to transparency are likely to continue shaping the development of artificial intelligence.

## See Also

- [large-language-model](https://www.wikiprompt.org/wiki/large-language-model)
- [artificial-intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence)
- [machine-learning](https://www.wikiprompt.org/wiki/machine-learning)
- [deep-learning](https://www.wikiprompt.org/wiki/deep-learning)
- [neural-network](https://www.wikiprompt.org/wiki/neural-network)
- [transformer](https://www.wikiprompt.org/wiki/transformer)
- [openai](https://www.wikiprompt.org/wiki/openai)
- [anthropic](https://www.wikiprompt.org/wiki/anthropic)
- [google-deepmind](https://www.wikiprompt.org/wiki/google-deepmind)
- [generative-ai](https://www.wikiprompt.org/wiki/generative-ai)

## References

- LMSYS official website: https://lmsys.org
- Chatbot Arena: https://chat.lmsys.org
- Vicuna model card on Hugging Face
- FastChat repository on GitHub

---
Source: https://www.wikiprompt.org/wiki/lm-sys
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-09T01:54:54.400758+00:00
