arXiv:2505.17512cs.AIcs.CL2025-05被引 1

用多人推理游戏测试大模型是否真懂概念,避免死记硬背。

Is Your LLM Really Mastering the Concept? A Multi-Agent Benchmark

  • 设计多人推理游戏,让AI通过对话辨析概念差异
  • 不同模型在概念理解上表现差异大,与整体能力不完全相关
  • 适合评估模型深层语义理解,尤其对安全与推理任务有参考价值

概念是支撑人类推理与分类的基本抽象。然而,大语言模型是否真正掌握这类概念结构,还是仅依赖表层模式记忆,尚不明确。现有基准多为静态、事实导向,难以探测细粒度语义理解,且易受数据泄露和过拟合影响。为此,我们提出CK-Arena,一个基于多智能体社交推理游戏(即《卧底》游戏)的动态概念知识评估框架。在该设定中,基于LLM的智能体被赋予细微不同的概念词,需通过描述、区分与推断他人陈述中的概念属性。模型性能通过游戏结果及生成描述的语义质量双重评估。此外,CK-Arena利用交互过程自动生成高质量问答数据,支持细粒度诊断分析。实验表明,概念理解在不同模型与类别间存在显著差异,且不严格对应总体模型能力。数据与代码已开源:https://ck-arena.site。

原文摘要 · Abstract (English)

Concepts serve as fundamental abstractions that support human reasoning and categorization. However, it remains unclear whether large language models truly capture such conceptual structures or primarily rely on surface-level pattern memorization. Existing benchmarks are largely static and fact oriented, which limits their ability to probe fine-grained semantic understanding and makes them vulnerable to data leakage and overfitting. To address this limitation, we introduce CK-Arena, a dynamic benchmark for conceptual knowledge evaluation based on a multi agent social deduction game, namely the Undercover game. In this setting, LLM based agents are assigned subtly different concept words and must describe, distinguish, and infer conceptual properties from others' statements. Model performance is evaluated through both game level outcomes and the semantic quality of generated descriptions. Furthermore, CK-Arena leverages the interaction process to automatically construct high quality question answering data for fine grained diagnostic analysis. Experimental results show that conceptual understanding varies substantially across models and categories, and is not strictly aligned with overall model capability. The data and code are available at the project homepage: https://ck-arena.site.

概念理解多智能体评测基准推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。