让AI理解多房间家居场景的结构关系,突破单间局限
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

- 用图神经网络增强物体上下文,引入房间级抽象令牌
- 通过分层注意力掩码实现基于场景拓扑的信息路由
- 在多房间任务上显著超越现有模型,适合复杂环境推理
现有3D场景基底大语言模型主要针对简化单间3D场景问答,难以处理包含多个连通房间和多样物体类别的真实家庭环境。我们提出CAIRN,一种拓扑感知的多房间3D场景理解3D-LLM。CAIRN将Transformer注意力与场景层级对齐,使模型显式具备物体级关系和房间级连通性的认知。它通过图神经网络为物体令牌注入房间内关系上下文,引入可学习的房间令牌进行房间级抽象,并采用带有几何偏置的分层注意力掩码,按场景拓扑路由信息。CAIRN基于我们在HM3D上构建的基准CAIRN-MR开发,涵盖定位、描述生成及四项逐步评估从室内感知到跨房间推理的任务。实验表明,CAIRN在所有CAIRN-MR任务上均大幅优于先前3D-LLM,同时在五个单间基准上保持竞争力。
原文摘要 · Abstract (English)
Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world household environments containing multiple interconnected rooms and diverse object categories. We introduce CAIRN, a topology-aware 3D-LLM for multi-room 3D scene understanding. CAIRN aligns transformer attention with scene hierarchy, giving the model explicit awareness of object-level relations and room-level connectivity. It enriches object tokens with room-local relational context via a graph neural network, introduces learned room tokens for room-level abstraction, and applies a hierarchical attention mask with geometric bias to route information according to scene topology. CAIRN is developed on CAIRN-MR, a benchmark we introduce on HM3D for multi-room 3D scene understanding, covering grounding, captioning, and four question-answering tasks that progressively evaluate from intra-room perception to cross-room reasoning. Experiments show that CAIRN outperforms prior 3D-LLMs by a large margin across all CAIRN-MR tasks while remaining competitive on five single-room benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。