受人眼中央与周边视觉启发,提升模型对3D场景的空间理解能力。
CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
- 用中央视觉类比的注意力聚焦机制,精准定位目标物体。
- 通过周边视觉类比的全局网格捕捉场景空间结构,提升推理效果。
- 适合需要复杂3D空间推理的机器人、自动驾驶任务。
我们提出一种受人眼中央-周边视觉启发的框架(CVP),这是一种简单而高效的多模态空间推理模型。现有方法主要依赖点云、体素或补丁特征等非结构化表示,并通过坐标嵌入隐式注入场景上下文,但常因缺乏显式的高层结构理解而导致空间推理能力有限。为此,我们在基于大模态模型的架构中引入两个互补组件:类中央视觉的目标亲和令牌,引导模型注意力聚焦于查询相关物体;类周边视觉的相对坐标网格,捕获全局场景上下文与空间布局。二者协同工作,实现对复杂3D环境的结构化、上下文感知理解。实验表明,CVP在多个3D场景理解基准上达到当前最优性能。
原文摘要 · Abstract (English)
We present a central-peripheral vision-inspired framework (CVP), a simple yet effective multimodal model for spatial reasoning that draws inspiration from the two types of human visual fields -- central vision and peripheral vision. Existing approaches primarily rely on unstructured representations, such as point clouds, voxels, or patch features, and inject scene context implicitly via coordinate embeddings. However, this often results in limited spatial reasoning capabilities due to the lack of explicit, high-level structural understanding. To address this limitation, we introduce two complementary components into a Large Multimodal Model-based architecture: target-affinity token, analogous to central vision, that guides the model's attention toward query-relevant objects; and allocentric grid, akin to peripheral vision, that captures global scene context and spatial arrangements. These components work in tandem to enable structured, context-aware understanding of complex 3D environments. Experiments show that CVP achieves state-of-the-art performance across a range of 3D scene understanding benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。