arXiv:2502.14359cs.CL2025-02EMNLP被引 14

用游戏和认知测试更有效区分大模型能力高低

Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests

  • 对比标准评测与互动游戏,发现游戏更能分辨模型优劣
  • 逻辑推理与工作记忆等能力与各类测试表现相关
  • 适合关注模型深层认知能力评估的研究者参考

我们考察了三种评估范式:标准基准(如 MMLU 和 BBH)、互动游戏(如信号博弈或禁忌词游戏)以及认知测试(如工作记忆或心智理论测试)。首先,探究在区分不同质量大模型方面,基准测试与互动游戏何者更有效。随后,受人类认知评估启发,我们整理了一套针对语言使用关键认知能力的专项测试,并研究其与模型在基准和游戏中的表现的相关性。分析显示,互动游戏在区分模型方面优于标准基准。因果与逻辑推理同时与静态和互动测试相关,而核心执行功能及社交/情绪能力则更多与游戏表现相关。我们主张开发受人类认知评估启发但专为大模型设计的新互动基准和靶向认知任务。

原文摘要 · Abstract (English)

We examine three evaluation paradigms: standard benchmarks (e.g., MMLU and BBH), interactive games (e.g., Signalling Games or Taboo), and cognitive tests (e.g., for working memory or theory of mind). First, we investigate which of the former two-benchmarks or games-is most effective at discriminating LLMs of varying quality. Then, inspired by human cognitive assessments, we compile a suite of targeted tests that measure cognitive abilities deemed essential for effective language use, and we investigate their correlation with model performance in benchmarks and games. Our analyses reveal that interactive games are superior to standard benchmarks in discriminating models. Causal and logical reasoning correlate with both static and interactive tests, while differences emerge regarding core executive functions and social/emotional skills, which correlate more with games. We advocate for the development of new interactive benchmarks and targeted cognitive tasks inspired by assessing human abilities but designed specifically for LLMs.

大模型评估认知测试互动游戏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。