用多人桌游测试视觉语言模型的推理能力,发现大模型更擅长猜图,小模型更会编谜。
DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
- 设计动态多角色游戏环境,评估模型生成和选择解释性假设的能力。
- 大模型在猜图任务中表现更好,小模型则更善于创造创意线索。
- 提出新评测基准,适合研究模型在模糊信息下的智能推理能力。
多模态溯因推理——从部分观测中生成并选择解释性假说——是智能的核心。当前对视觉语言模型(VLMs)的评估多局限于静态单角色任务。受桌游Dixit启发,我们提出DixitWorld,一个全面评估该能力的基准体系。DixitWorld包含两个核心组件:DixitArena,一个动态多代理环境,评估在不完全信息下“讲故事者”生成隐晦线索与“听众”从干扰项中选择目标图像的能力;DixitBench,一个静态问答基准,专门用于高效、可控地评估听众任务。DixitArena结果显示角色依赖行为:较小的开源模型通常更擅长创造性叙事,生成富有想象力但区分度较低的线索;较大的专有模型整体表现更优,尤其在听者任务中。DixitBench的表现与DixitArena中的听者结果高度相关,验证其作为假说选择可靠代理的有效性。研究揭示了多模态溯因推理中生成创造力与判别理解之间的关键权衡,这是构建更均衡、更强大视觉语言代理的核心挑战。
原文摘要 · Abstract (English)
Multimodal abductive reasoning--the generation and selection of explanatory hypotheses from partial observations--is a cornerstone of intelligence. Current evaluations of this ability in vision-language models (VLMs) are largely confined to static, single-agent tasks. Inspired by Dixit, we introduce DixitWorld, a comprehensive evaluation suite designed to deconstruct this challenge. DIXITWORLD features two core components: DixitArena, a dynamic, multi-agent environment that evaluates both hypothesis generation (a "storyteller" crafting cryptic clues) and hypothesis selection ("listeners" choosing the target image from decoys) under imperfect information; and DixitBench, a static QA benchmark that isolates the listener's task for efficient, controlled evaluation. Results from DixitArena reveal distinct, role-dependent behaviors: smaller open-source models often excel as creative storytellers, producing imaginative yet less discriminative clues, whereas larger proprietary models demonstrate superior overall performance, particularly as listeners. Performance on DixitBench strongly correlates with listener results in DixitArena, validating it as a reliable proxy for hypothesis selection. Our findings reveal a key trade-off between generative creativity and discriminative understanding in multimodal abductive reasoning, a central challenge for developing more balanced and capable vision-language agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。