用2D视觉标注提升4D场景图生成,解决数据少和词汇外问题。
Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene
- 通过4D大模型与3D掩码解码器端到端生成4D场景图。
- 链式推理机制利用大模型开放词汇能力,迭代补全对象与关系。
- 2D到4D特征迁移缓解数据稀缺,适合动态场景理解研究者。
最新提出的4D全景场景图(4D-PSG)为全面建模动态4D视觉世界提供了前所未有的表示。然而,现有4D-PSG研究面临严重数据稀缺问题,导致词汇外难题;同时基准生成流程的流水线特性也限制了性能。为此,本文提出一种新框架,利用丰富的2D视觉场景标注增强4D场景学习。首先,引入集成3D掩码解码器的4D大语言模型(4D-LLM),实现4D-PSG的端到端生成。进一步设计链式场景图推理机制,利用大模型的开放词汇能力,迭代推断准确且完整的对象与关系标签。最重要的是,提出2D到4D的视觉场景迁移学习框架,通过时空场景超越策略,将丰富2D场景图标注中的维度不变特征有效迁移到4D场景中,显著缓解4D-PSG的数据稀缺问题。在基准数据集上的大量实验表明,本方法显著优于基线模型,验证了其有效性。
原文摘要 · Abstract (English)
The latest emerged 4D Panoptic Scene Graph (4D-PSG) provides an advanced-ever representation for comprehensively modeling the dynamic 4D visual real world. Unfortunately, current pioneering 4D-PSG research can primarily suffer from data scarcity issues severely, as well as the resulting out-of-vocabulary problems; also, the pipeline nature of the benchmark generation method can lead to suboptimal performance. To address these challenges, this paper investigates a novel framework for 4D-PSG generation that leverages rich 2D visual scene annotations to enhance 4D scene learning. First, we introduce a 4D Large Language Model (4D-LLM) integrated with a 3D mask decoder for end-to-end generation of 4D-PSG. A chained SG inference mechanism is further designed to exploit LLMs' open-vocabulary capabilities to infer accurate and comprehensive object and relation labels iteratively. Most importantly, we propose a 2D-to-4D visual scene transfer learning framework, where a spatial-temporal scene transcending strategy effectively transfers dimension-invariant features from abundant 2D SG annotations to 4D scenes, effectively compensating for data scarcity in 4D-PSG. Extensive experiments on the benchmark data demonstrate that we strikingly outperform baseline models by a large margin, highlighting the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。