构建首个多模态模型组合类比推理诊断基准,揭示当前模型严重不足。
CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
- 设计多对组合类比任务,要求模型从多组图像中提取规则并组合
- 顶尖模型Gemini-2.5 Pro仅40.4%准确率,远低于人类100%
- 揭示模型在规则分解与复杂场景鲁棒性上的两大缺陷,适合研究者评估
类比推理测试人类认知的核心能力:将一对对象间的关系映射到另一对。现有对多模态大模型(MLLMs)类比能力的评估忽略了从多个来源组合规则的能力,而这是高阶智能的关键。为此,我们提出CARV(视觉中的组合类比推理),一个新任务及包含5,500个样本的数据集,作为首个诊断性基准。我们将类比从单对扩展到多对,要求MLLMs从每对中提取符号规则并组合生成新变换。对最先进MLLMs的评估显示显著性能差距:即使是最优的Gemini-2.5 Pro也仅达40.4%准确率,远低于人类水平的100%。诊断分析揭示两种持续失败模式:(1) 将视觉变化分解为符号规则,(2) 在多样或复杂设置下缺乏鲁棒性,凸显当前MLLMs在此任务上的根本局限。
原文摘要 · Abstract (English)
Analogical reasoning tests a fundamental aspect of human cognition: mapping the relation from one pair of objects to another. Existing evaluations of this ability in multimodal large language models (MLLMs) overlook the ability to compose rules from multiple sources, a critical component of higher-order intelligence. To close this gap, we introduce CARV (Compositional Analogical Reasoning in Vision), a novel task together with a 5,500-sample dataset as the first diagnostic benchmark. We extend the analogy from a single pair to multiple pairs, which requires MLLMs to extract symbolic rules from each pair and compose new transformations. Evaluation on the state-of-the-art MLLMs reveals a striking performance gap: even Gemini-2.5 Pro achieving only 40.4% accuracy, far below human-level performance of 100%. Diagnostic analysis shows two consistent failure modes: (1) decomposing visual changes into symbolic rules, and (2) maintaining robustness under diverse or complex settings, highlighting the limitations of current MLLMs on this task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。