开源通用大模型评估平台,支持多领域高效评测。
OpenCompass: A Universal Evaluation Platform for Large Language Models

- 模块化架构支持灵活配置与高并发评估
- 覆盖知识、推理、代码等多领域基准数据集
- 适合研究者与工业界进行大模型能力对比与优化
近年来,人工智能领域从任务专用的小规模模型转向通用大语言模型(LLMs)。随着大模型快速迭代,对其能力进行客观、量化、全面的评估成为技术发展的关键环节。当前主流静态基准数据集评估方法面临任务类型多样、评价标准不一、数据与流程碎片化等问题,难以实现跨领域、大规模模型的高效评估。为此,本文提出并开源了OpenCompass——一个一站式、可扩展、支持高并发的通用大模型评估平台。平台采用模块化与组件解耦设计,具备高兼容性、灵活性和高并发能力。核心架构包含五部分:配置系统、任务拆分模块、执行调度模块、任务执行单元与结果可视化模块。提供基于规则、大模型作为裁判及级联评估器等多种评估方式,适配不同任务场景。支持涵盖知识、推理、计算、科学、语言、代码等多个领域的主流基准数据集,为学术界与产业界提供统一高效的评估工具,助力精准识别大模型优劣势并推动后续优化。
原文摘要 · Abstract (English)
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。