让大模型像人一样看多张图推理空间关系
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
- 用两种目标显式训练模型理解跨视角对应区域
- 在三个基准上超越同类模型,接近更大模型表现
- 适合需要多视角空间推理的视觉问答任务
尽管多模态大语言模型在单图空间推理上取得进展,但需整合多视角信息的多图空间推理仍具挑战。认知研究表明,人类通过两种机制解决此类问题:跨视图对应(识别不同视角中对应同一物理位置的区域)和逐步视角变换(顺序组合相对视角变化)。现有研究仅部分且隐式引入这些机制,缺乏对两者的显式监督。我们提出人类感知训练框架HATCH,包含两个互补目标:(1) 图块级空间对齐,促使跨视角的空间对应区域图块表示对齐;(2) 动作后回答推理,要求模型在预测最终答案前生成明确的视角转换动作。在三个基准上的实验表明,HATCH在可比规模模型中持续显著优于基线,并达到与更大模型相当的效果,同时保持单图推理能力。
原文摘要 · Abstract (English)
While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, which requires integration of information from multiple viewpoints, remains challenging. Cognitive studies suggest that humans address such tasks through two mechanisms: cross-view correspondence, which identifies regions across different views that correspond to the same physical locations, and stepwise viewpoint transformation, which composes relative viewpoint changes sequentially. However, existing studies incorporate these mechanisms only partially and often implicitly, without explicit supervision for both. We propose Human-Aware Training for Cross-view correspondence and viewpoint cHange (HATCH), a training framework with two complementary objectives: (1) Patch-Level Spatial Alignment, which encourages patch representations to align across views for spatially corresponding regions, and (2) Action-then-Answer Reasoning, which requires the model to generate explicit viewpoint transition actions before predicting the final answer. Experiments on three benchmarks demonstrate that HATCH consistently outperforms baselines of comparable size by a clear margin and achieves competitive results against much larger models, while preserving single-image reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。