arXiv:2412.11314cs.CL2024-12中稿 · COLING 2025 system…被引 4

Evalica让模型评测更快更可靠,支持多方式接入。

Reliable, Reproducible, and Really Fast Leaderboards with Evalica

  • 开源工具包,支持网页、命令行、Python API三种接入方式
  • 提升评测的可重复性与可靠性,适用于LLM等NLP模型评估
  • 适合研究者和开发者快速搭建可信的模型排行榜

自然语言处理技术(如指令微调的大语言模型)的快速发展,亟需结合人类与机器反馈的现代评估协议。我们提出Evalica,一个开源工具包,用于构建可靠且可复现的模型排行榜。本文介绍其设计原理,评估其性能,并通过网页界面、命令行接口和Python API展示其可用性。

原文摘要 · Abstract (English)

The rapid advancement of natural language processing (NLP) technologies, such as instruction-tuned large language models (LLMs), urges the development of modern evaluation protocols with human and machine feedback. We introduce Evalica, an open-source toolkit that facilitates the creation of reliable and reproducible model leaderboards. This paper presents its design, evaluates its performance, and demonstrates its usability through its Web interface, command-line interface, and Python API.

模型评测开源工具大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。