arXiv:2602.02974cs.CV2026-02中稿 · as an IEEE TVCG pa…被引 2

从视频序列生成符合布局的3D场景,让混合现实内容适配用户空间。

SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequences

  • 用图网络+注意力机制预测语义场景图,捕捉物体间关系。
  • 在3RScan和SG-FRONT数据集上优于现有方法,复杂环境仍稳定生成。
  • 适合做空间化MR应用或需要精准布局生成的研究者。

我们提出SceneLinker,一种通过RGB序列生成组合式3D场景的新框架,利用语义场景图实现。为使混合现实(MR)内容自适应用户空间,需紧凑捕捉周围语义线索并还原真实布局。以往工作难以完整建模物体间的上下文关系,或仅关注形状多样性,导致生成场景与实际布置不符。为此,我们设计了具备交叉验证特征注意力的图网络用于场景图预测,并构建了包含联合形状与布局模块的图-变分自编码器(graph-VAE)以生成3D场景。在3RScan/3DSSG和SG-FRONT数据集上的实验表明,本方法在定量与定性评估中均超越当前最优模型,即使在复杂室内环境及严苛场景图约束下也表现优异。该工作使用户可基于物理环境生成一致的3D空间,从而创建空间化的MR内容。项目页面:https://scenelinker2026.github.io。

原文摘要 · Abstract (English)

We introduce SceneLinker, a novel framework that generates compositional 3D scenes via semantic scene graph from RGB sequences. To adaptively experience Mixed Reality (MR) content based on each user's space, it is essential to generate a 3D scene that reflects the real-world layout by compactly capturing the semantic cues of the surroundings. Prior works struggled to fully capture the contextual relationship between objects or mainly focused on synthesizing diverse shapes, making it challenging to generate 3D scenes aligned with object arrangements. We address these challenges by designing a graph network with cross-check feature attention for scene graph prediction and constructing a graph-variational autoencoder (graph-VAE), which consists of a joint shape and layout block for 3D scene generation. Experiments on the 3RScan/3DSSG and SG-FRONT datasets demonstrate that our approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations, even in complex indoor environments and under challenging scene graph constraints. Our work enables users to generate consistent 3D spaces from their physical environments via scene graphs, allowing them to create spatial MR content. Project page is https://scenelinker2026.github.io.

3D生成场景图MR视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。