用博弈竞争方式评估大模型,让模型在对抗中展现真实能力。
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
- 设计多种对抗性游戏评估模型能力
- 覆盖策略推理、知识应用等多维度表现
- 支持扩展且易实现,适合研究者快速测试
我们提出 ZeroSumEval,一个基于竞争的动态评估框架,用于评测大语言模型(LLM)。该框架通过安全挑战(如 Capture the Flag)、经典棋类(如 chess)和知识测试(如 MathQuiz)等多样化游戏,评估模型的战略推理、规划、知识应用、安全性与适应性等能力。基于近期研究表明游戏化评估对 LLM 有效,ZeroSumEval 进一步提供标准化、可扩展的游戏实现方式,并利用 DSPy 实现更优的模型策略抽象,便于研究人员高效构建与评估。
原文摘要 · Abstract (English)
We introduce ZeroSumEval, a dynamic, competition-based, and evolving evaluation framework for Large Language Models (LLMs) that leverages competitive games. ZeroSumEval encompasses a diverse suite of games, including security challenges (Capture the Flag), classic board games (chess), and knowledge tests (MathQuiz). These games are designed to evaluate a range of capabilities such as strategic reasoning, planning, knowledge application, safety, and adaptability. Building upon recent studies that highlight the effectiveness of game-based evaluations for LLMs, ZeroSumEval enhances these approaches by providing a standardized and extensible framework for easily implementing games and leverages DSPy to provide a better abstraction for LLM player strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。