arXiv:2510.19892cs.CLcs.AI2025-10中稿 · EMNLP被引 1

用游戏Dixit评估多模态大模型综合能力,更真实可靠。

Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities

  • 以卡牌游戏Dixit为评测场景,综合考察模型生成与推理能力。
  • 五款模型在游戏中的胜率与主流基准排名完全一致。
  • 揭示模型策略缺陷,适合研究多模态推理与人类差异的学者。

多模态大模型通常在静态、独立的基准上评估,难以全面检验其综合能力;或依赖人工或模型间的两两比较,主观性强且易被表面特征(如冗长)误导。为此,我们提出基于游戏的评估方法,利用游戏具备多重能力需求、竞争性及明确规则的特点,构建更客观、全面的评测框架。具体通过奇幻卡牌游戏Dixit实现:玩家需为一张卡片生成能误导部分人但不误导所有人的描述。对五款MLM的定量实验显示,其在Dixit中的胜率排名与主流基准完全一致;而人机对战分析揭示了模型策略差异与改进空间,验证了该方法的有效性。

原文摘要 · Abstract (English)

Multi-modal large language models (MLMs) are often assessed on static, individual benchmarks -- which cannot jointly assess MLM capabilities in a single task -- or rely on human or model pairwise comparisons -- which is highly subjective, expensive, and allows models to exploit superficial shortcuts (e.g., verbosity) to inflate their win-rates. To overcome these issues, we propose game-based evaluations to holistically assess MLM capabilities. Games require multiple abilities for players to win, are inherently competitive, and are governed by fix, objective rules, and makes evaluation more engaging, providing a robust framework to address the aforementioned challenges. We manifest this evaluation specifically through Dixit, a fantasy card game where players must generate captions for a card that trick some, but not all players, into selecting the played card. Our quantitative experiments with five MLMs show Dixit win-rate rankings are perfectly correlated with those on popular MLM benchmarks, while games between human and MLM players in Dixit reveal several differences between agent strategies and areas of improvement for MLM reasoning.

多模态游戏评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。