用对话游戏评测大模型,兼顾可控性与真实场景测试。
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
- 设计对话游戏框架,实现可重复、无参考的多轮交互评估。
- 通过clembench平台支持自定义任务扩展,适配不同模型评测需求。
- 适合希望在真实交互中系统评估大模型能力的研究者或开发者。
目前大语言模型评估主要有两种范式:基于参考答案的评估和基于用户偏好的评估。前者控制性强但缺乏真实场景,后者生态效度高但难以重复。近年来兴起第三种范式——对话游戏评估,结合两者优势,强调目标导向的多轮交互,兼具可控性与真实性。然而其推广受限于缺乏成熟易用的实现工具。本文提出clembench,自2023年起持续开发,最新版本优化了通用性与易用性。它支持使用预设英文对话游戏实例对自有模型进行基准测试,并可轻松扩展新定制化测试任务,推动对话游戏评估的普及应用。
原文摘要 · Abstract (English)
There are currently two main paradigms for evaluating large language models (LLMs), reference-based evaluation and preference-based evaluation. The first, carried over from the evaluation of machine learning models in general, relies on pre-defined task instances, for which reference task executions are available. The second, best exemplified by the LM-arena, relies on (often self-selected) users bringing their own intents to a site that routes these to several models in parallel, among whose responses the user then selects their most preferred one. The former paradigm hence excels at control over what is tested, while the latter comes with higher ecological validity, testing actual use cases interactively. Recently, a third complementary paradigm has emerged that combines some of the strengths of these approaches, offering control over multi-turn, reference-free, repeatable interactions, while stressing goal-directedness: dialogue game based evaluation. While the utility of this approach has been shown by several projects, its adoption has been held back by the lack of a mature, easily re-usable implementation. In this paper, we present clembench, which has been in continuous development since 2023 and has in its latest release been optimized for ease of general use. We describe how it can be used to benchmark one's own models (using a provided set of benchmark game instances in English), as well as how easily the benchmark itself can be extended with new, tailor-made targeted tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。