评测大模型多模态关系推理能力,发现错误主因是推理与记忆不足。
SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction

- 构建自适应对话基准SciReC,覆盖多种关系推理任务。
- Claude 4.6表现最佳(73%),开源模型在空间关系上最弱。
- 诊断框架揭示推理与记忆是主要失败原因,适合模型优化研究者。
关系推理涉及对概念间潜在关系的感知理解、比较与整合,涵盖类比、结构和因果等多种类型,反映高层次认知能力。为评估多模态大语言模型(MLLM)在这些任务上的表现,我们提出了SciReC——一个模型自适应的多模态学术对话基准。由于关系推理涉及多重表征与多种因素(视觉理解、知识展现、记忆回忆),我们设计了DMRA缺陷诊断框架,量化各组件贡献以定位失败主因。实验显示,Claude 4.6在整体关系得分上最高(73%),紧随其后的是GPT 5.4(68%)。性能趋势表明,开源模型在空间关系任务上得分最低,而专有模型更难处理层级与序列关系。跨领域表现中,天文学领域得分最低,心理学最高。DMRA分析结果表明,关系推理能力不足是所有模型的主要错误来源,其次为记忆限制。
原文摘要 · Abstract (English)
Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts. This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different aspect of higher-order understanding. To examine the performance of multimodal large language models (MLLM) on these relational inference tasks, we developed SciReC, a model-adaptive multimodal academic dialog benchmark. As the relational reasoning process involves multiple representations and various factors (visual understanding, exhibiting knowledge, and memory recall), we propose DMRA, a deficit-based diagnostic framework that quantifies the contribution of these components to identify the primary cause of unsuccessful cases. Claude 4.6 achieved the best performance on the overall relational score with 73\%, followed by GPT 5.4 with 68\%. Performance trends indicate that open-source models achieve their lowest scores on spatial relations, while proprietary models struggle more with hierarchical and sequential relations. Across domains, model performance is lowest on Astronomy and highest on Psychology. The results of DMRA reveal that relational reasoning is the primary source of error across all models, followed by memory limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。