← Open Source
open-compass

opencompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100+ datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.

AI EngineeringTest & guardPython
Open on GitHub
Momentum
+4stars in 24 hours+0.1%
7.51k
Stars
882
Forks
+16
This week
100
Contributors
Created 2023-06-15 · Updated 2026-10-10 · #1997 today
Top developers
README

🌐Website | 📖CompassHub | 📊CompassRank | 📘Documentation | 🛠️Installation | 🤔Reporting Issues

English | 简体中文

👋 join us on [Discord](https://discord.gg/KKwfEbFj7U) and [WeChat](https://r.vansin.top/?r=opencompass)

[!IMPORTANT]

Star Us, You will receive all release notifications from GitHub without any delay ~ ⭐️

Star History

🧭 Welcome

to OpenCompass!

Just like a compass guides us on our journey, OpenCompass will guide you through the complex landscape of evaluating large language models. With its powerful algorithms and intuitive interface, OpenCompass makes it easy to assess the quality and effectiveness of your NLP models.

🚩🚩🚩 Explore opportunities at OpenCompass! We're currently hiring full-time researchers/engineers and interns. If you're passionate about LLM and OpenCompass, don't hesitate to reach out to us via email. We'd love to hear from you!

🔥🔥🔥 We are delighted to announce that the OpenCompass has been recommended by the Meta AI, click Get Started of Llama for more information.

✨ Introduction

image

OpenCompass is a one-stop platform for large model evaluation. It supports a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, and more, and integrates over 100 datasets covering knowledge, reasoning, coding, science, language, long context, safety, and other capability dimensions. OpenCompass also integrates VLMEvalKit, enabling text-only and vision-language datasets to be evaluated within a unified workflow.

OpenCompass provides a complete workflow spanning dataset and model configuration, inference, evaluation, and result summarization. It supports local open-source models and API models, as well as multiple inference backends such as HuggingFace, LMDeploy, and vLLM. The platform supports zero-shot, few-shot, chain-of-thought, rule-based, and LLM-as-judge evaluation, while task partitioning, concurrent execution, and distributed execution accommodate evaluation workloads of different scales.

OpenCompass is designed to be open, reproducible, and extensible. Users can flexibly integrate new models, datasets, evaluators, inference backends, and task scheduling systems. The accompanying CompassHub provides benchmark navigation, while CompassRank presents model evaluation leaderboards.

🚀 What's New

  • [2026.09.21] OpenCompass has updated the recommended API configurations for flagship models from leading providers and comprehensively restructured its Chinese and English documentation, further improving the model integration and evaluation experience. See the model configurations and OpenCompass documentation for details! 🔥🔥🔥
  • [2026.08.25] OpenCompass now integrates with VLMEvalKit, enabling native multimodal dataset loading, inference through OpenAI-compatible APIs, and evaluation with official VLMEvalKit metrics. Check out the MMBench example and MMMU-Pro example for details! 🔥🔥🔥
  • [2026.07.28] OpenCompass has expanded its API model ecosystem with support for the OpenAI Responses API and LiteLLM AI Gateway, while updating the Gemini and Anthropic integrations to their latest SDK interfaces. Check out the OpenAI Responses API implementation, LiteLLM AI Gateway implementation, Gemini SDK implementation, and Anthropic SDK implementation for details!
  • [2026.07.27] OpenCompass now supports multi-round inference in GenInferencer and adds support for the Multi-IF dataset to evaluate multi-turn instruction-following capabilities. Check out the Multi-IF evaluation configuration for details! 🔥🔥🔥
  • [2026.05.25] OpenCompass now provides repeat analysis tools for detecting repetitive content and looping model outputs in current evaluation tasks or existing evaluation results. Check out the repeat analysis tool for details!
  • [2026.03.20] OpenCompass now supports concurrent inference across tasks together with evaluation watching, enabling completed inference tasks to be monitored and subsequent evaluations to be triggered in a coordinated pipeline. Parallel inferencers, task monitoring, and heartbeat mechanisms further improve large-scale evaluation efficiency. Check out the concurrent inference implementation and evaluation watcher implementation for details!
  • [2026.03.17] OpenCompass introduces RawPromptTemplate, allowing original benchmark prompts and structured conversations to be passed to models without unintended formatting transformations. It supports API models, ChatML datasets, and appending additional prompt content on the model side. Check out the RawPromptTemplate guide for details!
  • [2026.02.05] OpenCompass now supports Intern-S1-Pro related general and scientific evaluation benchmarks. Please check Example for Evaluating Intern-S1-Pro and Model Card for more details! 🔥🔥🔥
  • [2025.12.08] OpenCompass now supports evaluation for SciReasoner. Please check Example for Evaluating SciReasoner and Project GitHub Repo for more details! 🔥🔥🔥
  • [2025.07.26] OpenCompass now supports Intern-S1 related general and scientific evaluation benchmarks. Please refer to the Intern-S1 model configuration for details! 🔥🔥🔥
  • [2025.04.01] OpenCompass now supports CascadeEvaluator, allowing multiple evaluators to work in sequence and enabling custom evaluation pipelines for more complex scenarios. Check out the documentation for details! 🔥🔥🔥
  • [2025.03.11] OpenCompass now supports SuperGPQA, covering knowledge evaluation across 285 graduate-level disciplines. Give it a try! 🔥🔥🔥
  • [2025.02.28] OpenCompass now supports the DeepSeek-R1 model series. Check out the DeepSeek-R1 model configuration for more details! 🔥🔥🔥
  • [2025.02.15] We have added two practical evaluation tools: GenericLLMEvaluator for LLM-as-judge evaluation and MATHVerifyEvaluator for mathematical reasoning evaluation. Check out the LLM Judge and Mathematical Evaluation documentation for more details! 🔥🔥🔥

🛠️ Installation

Below are the steps for quick installation and dataset preparation.

💻 Environment Setup

We highly recommend using conda to manage your Python environment. OpenCompass supports Python 3.12 for regular and full installations. If your evaluation depends on code execution datasets backed by pyext, use Python 3.10 instead: pyext==0.7 is skipped on Python >=3.11 because it relies on inspect.getargspec, which was removed in Python 3.11. Without pyext, APPS (apps, apps_mini), TACO, and LiveCodeBench Code Generation are unavailable.

  • Create your virtual environment

    conda create --name opencompass python=3.12 -y
    conda activate opencompass
    
  • Install OpenCompass via pip

    # Supports most datasets and models
    pip install -U opencompass
    
    # Full installation (supports more datasets)
    # pip install "opencompass[full]"
    
    # Model inference backends. Since these backends often have conflicting dependencies,
    # we recommend managing them in separate virtual environments.
    # pip install "opencompass[lmdeploy]"
    # pip install "opencompass[vllm]"
    
    # API evaluation (for example, OpenAI and Qwen)
    # pip install "opencompass[api]"
    
    # Multimodal evaluation
    # pip install "opencompass[vlm]"
    
  • Install OpenCompass from source

    To use the latest OpenCompass features, you can also build it from source:

    git clone https://github.com/open-compass/opencompass opencompass
    cd opencompass
    pip install -e .
    # pip install -e ".[full]"
    # pip install -e ".[vllm]"
    

After installation, read the Quick Start to learn how to run an evaluation task.

For more tutorials, see our full documentation.

🔝Back to top

📖 Dataset Support

The OpenCompass documentation provides a statistical list of all datasets supported by the platform.

You can quickly find the dataset you need by sorting, filtering, and searching the list.

See the dataset statistics section of the dataset documentation for details.

🔝Back to top

📖 Model Support

    Open-source Models
  

  

    API Models

🔝Back to top

📊 Leaderboard

We will continue to provide detailed leaderboards for open-source and API models. See the OpenCompass Leaderboard. To participate in an evaluation, send the model repository URL or a standard API endpoint via email.

🔝Back to top

👷‍♂️ Contributing

We appreciate all contributions to improving OpenCompass. Please refer to the contributing guideline for the best practice.

[

](https://github.com/open-compass/opencompass/graphs/contributors)

🤝 Acknowledgements

Some code in this project is cited and modified from OpenICL.

Some datasets and prompt implementations are modified from chain-of-thought-hub and instruct-eval.

🖊️ Citation

@misc{2023opencompass,
    title={OpenCompass: A Universal Evaluation Platform for Foundation Models},
    author={OpenCompass Contributors},
    howpublished = {\url{https://github.com/open-compass/opencompass}},
    year={2023}
}

🔝Back to top