arXiv:2606.03577cs.CV2026-06被引 1

用新基准和强化学习提升大模型的空间推理能力

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

  • 设计分层基准测试,评估视角变化下的空间匹配能力
  • 在90样本难题上,人类F1达84.0,现有模型仅37.2
  • 自动生成数据并用可验证奖励训练,无需复杂提示

宽基线匹配(WBM)需融合几何理解、视角变换、细粒度感知与遮挡推理,是检验多模态大模型在真实环境中的空间推理能力的挑战性任务。然而当前模型缺乏系统性评估与训练框架。本文提出ReasonMatch-Bench,按视角偏移和匹配粒度划分,覆盖室内、室外及物体中心场景;在90个高难度样本子集上,人工标注者达到84.0的F1分数,而最优基线仅37.2。为此,我们构建可扩展的数据生成流水线,从大规模视频-3D语料库(包括RGB-D视频与SfM重建)中自动提取宽基线视图对,提供多样且可验证的监督信号。进一步提出动态对应强化学习(DCRL),结合图像级视角演进与点级对应课程学习,通过可验证奖励实现无显式思维链监督的训练。大量实验表明,DCRL显著提升ReasonMatch-Bench表现,并迁移至相关空间基准,同时在多个通用视觉基准上保持性能,仅有小幅提升。

原文摘要 · Abstract (English)

Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making it a challenging testbed for spatial reasoning in multimodal large language models (MLLMs) deployed in physical environments. However, current MLLMs lack systematic evaluation and training frameworks for these capabilities. We introduce ReasonMatch-Bench, a benchmark stratified by viewpoint displacement and matching granularity across indoor, outdoor, and object-centric scenarios, and show that current MLLMs still struggle with fine-grained wide-baseline correspondence: on a difficult 90-sample subset, human annotators achieve 84.0 F1, while the best existing baseline reaches 37.2. To bridge this gap, we build a scalable data-generation pipeline that automatically extracts wide-baseline view pairs from large-scale video-3D corpora, including RGB-D videos and SfM reconstructions, yielding diverse and verifiable supervision. We further propose Dynamic Correspondence Reinforcement Learning (DCRL), which combines Image-Level Viewpoint Progression and Point-Level Correspondence Curriculum to improve WBM training through verifiable rewards without explicit CoT supervision. Extensive experiments show that DCRL substantially improves ReasonMatch-Bench and transfers to related spatial benchmarks, while maintaining general visual understanding performance with modest gains on several benchmarks.

空间推理多模态模型强化学习基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。