Evalica让模型评测更快更可靠,支持多方式接入。
Reliable, Reproducible, and Really Fast Leaderboards with Evalica
- 开源工具包,支持网页、命令行、Python API三种接入方式
- 提升评测的可重复性与可靠性,适用于LLM等NLP模型评估
- 适合研究者和开发者快速搭建可信的模型排行榜
自然语言处理技术(如指令微调的大语言模型)的快速发展,亟需结合人类与机器反馈的现代评估协议。我们提出Evalica,一个开源工具包,用于构建可靠且可复现的模型排行榜。本文介绍其设计原理,评估其性能,并通过网页界面、命令行接口和Python API展示其可用性。
原文摘要 · Abstract (English)
The rapid advancement of natural language processing (NLP) technologies, such as instruction-tuned large language models (LLMs), urges the development of modern evaluation protocols with human and machine feedback. We introduce Evalica, an open-source toolkit that facilitates the creation of reliable and reproducible model leaderboards. This paper presents its design, evaluates its performance, and demonstrates its usability through its Web interface, command-line interface, and Python API.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。