arXiv:2512.14364cs.CV2025-12中稿 · TMLR被引 6

一个模型搞定3D场景所有语义理解任务,仅用RGB图秒级输出完整语义几何。

Unified Semantic Transformer for 3D Scene Understanding

  • 统一框架整合分割、实例嵌入、开放词汇特征等多任务,端到端训练。
  • 仅需数秒完成全场景推理,性能超越多个专用模型,甚至优于依赖真值3D几何的方法。
  • 创新多视角一致性损失,结合自监督与2D蒸馏,适用于未知场景泛化。

整体化3D场景理解需要捕捉并解析非结构化的3D环境。由于现实世界的固有复杂性,现有模型大多为特定任务设计且功能受限。我们提出UNITE:一种用于3D场景理解的统一语义变压器,这是一种新型前馈神经网络,可在一个模型中统一多种3D密集语义室内任务。该模型在未见过的场景上以完全端到端方式训练,仅需数秒即可推断出完整的3D语义几何。方法仅需输入RGB图像,直接预测多项密集语义属性,包括3D场景分割、实例嵌入、开放词汇特征和关节结构。训练结合2D知识蒸馏,高度依赖自监督,并引入新颖的多视角损失以保证3D视角一致性。实验表明,UNITE在多个不同密集室内语义任务上达到当前最优性能,许多情况下甚至超越专门优化的任务模型,部分表现优于依赖真实3D几何信息的方法。

原文摘要 · Abstract (English)

Holistic 3D scene understanding involves capturing and parsing unstructured 3D environments. Due to the inherent complexity of the real world, existing models have predominantly been developed and limited to be task-specific. We introduce UNITE, a Unified Semantic Transformer for 3D scene understanding, a novel feed-forward neural network that unifies a diverse set of 3D dense semantic indoor tasks within a single model. Our model operates on unseen scenes trained in a fully end-to-end manner and only takes a couple seconds to infer the full 3D semantic geometry. Our approach is capable of directly predicting multiple dense semantic attributes, including 3D scene segmentation, instance embeddings, open-vocabulary features, and articulations, solely from RGB images. The method is trained using a combination of 2D distillation, heavily relying on self-supervision and leverages novel multi-view losses designed to ensure 3D view consistency. We demonstrate that UNITE achieves state-of-the-art performance on several different dense indoor semantic tasks and even outperforms task-specific models, in many cases, surpassing methods that operate on ground truth 3D geometry. See the project website at unite-page.github.io

3D理解统一模型语义分割自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。