arXiv:2607.17657cs.AIcs.CV2026-07

让多模态模型摆脱摄像头视角依赖,学会从物体自身方向推理空间关系。

OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment

论文配图:OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment
图 1 · 摘自论文原文
  • 通过方位感知标记和傅里叶角度编码显式建模物体朝向
  • 在非相机视角任务上性能显著提升,尤其在人物中心和朝向敏感任务中
  • 适用于需要精准空间推理的视觉语言模型,如机器人导航、人机交互

多模态大模型在需要视角转换的空间推理任务中仍表现不佳,常依赖摄像头视角而非参考物体的自身视角,导致非相机参考场景下系统性错误。本文分析该问题根源,发现物体朝向是导致摄像头视角捷径行为的关键因素。为此提出OrientSAM框架,通过方位感知标记与基于傅里叶的角度编码,将显式朝向信息注入多模态表示,并采用课程学习策略逐步提升视角感知推理能力。同时构建大规模图像生成的方位感知空间监督数据构造流水线。在Spatial-MM、ViewSpatial和3DSRBench上的实验表明,OrientSAM在各类非相机视角、人物中心及朝向敏感任务中持续优于强基线模型。结果进一步证明,显式建模朝向对缓解摄像头视角捷径、实现更鲁棒的非参照系空间推理至关重要。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) still struggle with spatial reasoning that requires perspective transformation. In particular, they often rely on camera-centric cues rather than reasoning from the reference object's viewpoint, leading to systematic errors in non-camera reference settings. In this paper, we first analyze this failure mode and show that object orientation is a key factor underlying such camera-centric shortcut behavior. To address this issue, we propose OrientSAM, an orientation-aware spatial alignment framework for multimodal models. OrientSAM injects explicit orientation information into multimodal representations through orientation-aware tokens and Fourier-based angle encoding, and further adopts a curriculum learning strategy to progressively improve perspective-aware reasoning. In addition, we build a spatial data construction pipeline to generate orientation-aware spatial supervision from large-scale images. Experiments on Spatial-MM, ViewSpatial, and 3DSRBench show that OrientSAM consistently outperforms strong baselines, especially on non-camera-view, person-centric, and orientation-sensitive tasks. The results further demonstrate that explicit orientation modeling is important for mitigating camera-centric shortcut behavior and enabling more robust allocentric spatial reasoning in multimodal models.

空间推理多模态视角不变方位感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。