让视觉语言模型理解2D图像与3D空间的区域关系,支持灵活标注。
3D Aware Region Prompted Vision Language Model
- 用3D位置嵌入增强2D特征,实现跨视角空间推理。
- 在2D和3D基准上均达到顶尖性能,支持非同步物体识别。
- 适用于真实视频场景,无需3D标注仍可准确推断空间关系。
我们提出一种空间区域3D(SR-3D)感知的视觉语言模型,通过共享的视觉标记空间连接单视图2D图像与多视图3D数据。SR-3D支持灵活的区域提示,用户可在任意帧上用边界框、分割掩码或直接在3D中标注,无需对所有帧进行详尽标注。通过将3D位置嵌入融入2D视觉特征,使3D模型能利用强2D先验,在对象未共现于同一视角时仍实现精确的跨帧空间推理。在通用2D视觉语言与专用3D空间基准上的大量实验表明,SR-3D实现了当前最优性能,验证了其在统一2D与3D表示空间方面的有效性。此外,我们发现该模型在无传感器3D输入或真实3D标注的野外视频中,仍能准确推断空间关系与度量信息。
原文摘要 · Abstract (English)
We present Spatial Region 3D (SR-3D) aware vision-language model that connects single-view 2D images and multi-view 3D data through a shared visual token space. SR-3D supports flexible region prompting, allowing users to annotate regions with bounding boxes, segmentation masks on any frame, or directly in 3D, without the need for exhaustive multi-frame labeling. We achieve this by enriching 2D visual features with 3D positional embeddings, which allows the 3D model to draw upon strong 2D priors for more accurate spatial reasoning across frames, even when objects of interest do not co-occur within the same view. Extensive experiments on both general 2D vision language and specialized 3D spatial benchmarks demonstrate that SR-3D achieves state-of-the-art performance, underscoring its effectiveness for unifying 2D and 3D representation space on scene understanding. Moreover, we observe applicability to in-the-wild videos without sensory 3D inputs or ground-truth 3D annotations, where SR-3D accurately infers spatial relationships and metric measurements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。