arXiv:2605.15708cs.CV2026-05

让3D模型理解'左'、'右'等视角相关空间关系,提升语义分割准确率。

3D Segmentation Using Viewpoint-Dependent Spatial Relationships

论文配图:3D Segmentation Using Viewpoint-Dependent Spatial Relationships
图 1 · 摘自论文原文
  • 构建220万样本的视角感知数据集,支持密集视角采样。
  • 引入视角编码机制后,模型在视角依赖关系上的平均交并比提升至0.47。
  • 适合研究多模态3D理解、视觉-语言对齐的学者使用。

近年来,3D数据集和多模态模型的发展显著提升了自然语言驱动的3D场景理解能力。然而,大多数3D指代分割方法未显式建模观察者视角,导致'左'、'右'、'前'、'后'等空间关系模糊且难以评估。本文提出一个视角感知的3D指代分割数据集,包含220k基准样本,可通过密集视角采样扩展至数千万样本。该数据集中目标物体仅能通过以观察者为中心的空间关系定位,因此必须进行视角条件化定位。我们利用相机位姿自动标注以观察者为中心的关系(左/右、前/后)以及与视角无关的关系(上/下)。基于此基准,我们在零样本设置下评估多个现有3D大模型,发现当前模型在视角依赖指令上表现不佳。进一步研究如何将显式视角信息融入3D大模型,提出一种编码相机位姿的视角表示,并以观察视角条件化模型,使分割准确率在视角依赖关系上提升,mIoU从0.30提高到0.47。数据集、代码及训练模型将在论文接受后公开。

原文摘要 · Abstract (English)

Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentation methods do not explicitly represent the observer viewpoint, making spatial relations such as "left," "right," "front," and "behind" ambiguous and difficult to evaluate. We introduce a viewpoint-aware 3D referring segmentation dataset containing 220k benchmark samples, and scalable to tens of millions of viewpoint-conditioned samples through dense viewpoint sampling. In this dataset, target objects can only be identified through observer-centric spatial relations, making viewpoint-conditioned grounding necessary. We construct the benchmark by leveraging camera poses to automatically annotate observer-centric relations (left/right, front/behind) together with viewpoint-independent relations (above/under). Using this benchmark, we evaluate several existing 3D large multimodal models in a zero-shot setting and find that current models struggle with viewpoint-dependent spatial instructions. We further study how explicit viewpoint information can be incorporated into 3D large multimodal models. We introduce a viewpoint representation that encodes camera poses and conditions the model on the observation viewpoint, improving segmentation accuracy on viewpoint-dependent relations and increasing mIoU from 0.30 to 0.47 compared to a model without viewpoint conditioning. The dataset, code, and trained models will be made publicly available upon acceptance.

3D分割空间关系多模态视角感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。