用扩散模型统一生成3D场景图的结构与特征,支持任意层级构建。
Generation of High-Level Concepts in 3D Scene Graphs via Autoregressive Diffusion

- 基于自回归扩散模型,端到端联合生成场景图结构与节点特征。
- 在合成、建筑图、机器人数据上均优于现有方法,最高提升18.7%。
- 适用于新概念类和复杂层次结构,适合机器人空间感知研究者。
室内3D场景图(3DSGs)将环境表示为多层层级结构,连接观测到的几何基元(如平面)与高层度量-语义概念(如房间、楼层、建筑),支持机器人感知与SLAM中的增量式空间推理。然而,传统高层概念生成方法依赖特定类别的人工规则,学习型方法则需分别建模图结构与空间节点特征(如质心),限制了对新类别和更复杂层级的扩展性。本文提出一种统一的自回归扩散图生成模型,从任意深度的垂直平面出发,自底向上联合学习结构与特征,构建完整的3DSGs。该方法在涵盖合成场景、真实建筑平面图及机器人传感器数据的多个3DSG数据集上,始终优于所有学习型与随机基线,在最大层级和真实单层数据上甚至超越已知目标图大小的一次性模型。最后,我们提出融合Gromov-Wasserstein距离的改进评估指标,实现对生成3DSGs与真实图之间的图级一致性评估。
原文摘要 · Abstract (English)
Indoor 3D Scene Graphs (3DSGs) represent environments as multi-layer hierarchies that connect observed geometric primitives (e.g., planes) to higher-level metric-semantic concepts (e.g., rooms, floors, buildings), enabling incremental spatial reasoning for robotic perception and SLAM. However, classical high-level concept generation approaches rely on hand-crafted rules for specific concept classes, while learning-based methods require separate models for graph structure and spatial node features (e.g., centroids), which limits scalability to novel classes and more complex hierarchies. We propose a unified autoregressive diffusion-based graph generative model that jointly learns structure and features, constructing complete 3DSGs bottom-up from observed vertical planes across arbitrary hierarchy depths. Our method consistently surpasses all learning-based and random baselines across 3DSG datasets spanning synthetic scenes, real architectural floor plans, and robotic sensor data, with varying layout complexity and hierarchy depth, and surpasses a one-shot model with oracle access to the target graph size on the largest hierarchy and on real single-floor data. Finally, we propose an adaptation of the Fused Gromov--Wasserstein distance for principled graph-level evaluation of generated 3DSGs against ground truth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。