arXiv:2412.06394cs.AIcs.CL2024-12ICLR被引 28

用真人互动游戏评估大模型推理能力,更真实也更精细。

GameArena: Evaluating LLM Reasoning through Live Computer Games

  • 设计三款互动游戏,分别测试演绎与归纳等推理能力
  • 收集2000+场次游戏数据,评估5个顶尖大模型的推理表现
  • 首次在真实环境中获取大模型的逐步推理过程,适合研究者使用

评估大语言模型(LLMs)的推理能力面临挑战。现有基准多依赖静态数据集,易受数据污染且会随时间饱和,或依赖二元化的人类实时反馈,混淆推理与其他能力。作为最突出的动态基准,Chatbot Arena 在真实场景中评估开放式问题,但缺乏对具体推理能力的细粒度分析。我们提出 GameArena,一个通过人机互动游戏评估 LLM 推理能力的动态基准。GameArena 包含三款游戏,分别测试特定推理能力(如演绎与归纳推理),同时保持参与者的兴趣与投入。我们回溯分析游戏数据,揭示 LLM 的内在推理过程,并量化其细粒度推理表现。共收集超过2000场游戏会话,对五个先进 LLM 提供详细推理能力评估。100名参与者用户研究显示,GameArena 比 Chatbot Arena 更能提升用户参与度。首次实现真实环境中对 LLM 推理过程的逐步数据采集。

原文摘要 · Abstract (English)

Evaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination and may get saturated over time, or on binary live human feedback that conflates reasoning with other abilities. As the most prominent dynamic benchmark, Chatbot Arena evaluates open-ended questions in real-world settings, but lacks the granularity in assessing specific reasoning capabilities. We introduce GameArena, a dynamic benchmark designed to evaluate LLM reasoning capabilities through interactive gameplay with humans. GameArena consists of three games designed to test specific reasoning capabilities (e.g., deductive and inductive reasoning), while keeping participants entertained and engaged. We analyze the gaming data retrospectively to uncover the underlying reasoning processes of LLMs and measure their fine-grained reasoning capabilities. We collect over 2000 game sessions and provide detailed assessments of various reasoning capabilities for five state-of-the-art LLMs. Our user study with 100 participants suggests that GameArena improves user engagement compared to Chatbot Arena. For the first time, GameArena enables the collection of step-by-step LLM reasoning data in the wild.

大模型评测推理能力交互游戏动态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。