arXiv:2512.10957cs.CVcs.AI2025-12被引 13

解耦去遮挡与姿态估计,实现复杂遮挡下的高质量3D场景生成

SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model

  • 将去遮挡与3D生成分离建模,提升开放集遮挡适应性
  • 融合全局与局部注意力机制,显著提升姿态估计精度
  • 构建新数据集,支持开放集场景泛化,适合3D生成研究者

本文提出一种解耦式3D场景生成框架SceneMaker。由于缺乏充分的开放集去遮挡与姿态估计先验,现有方法在严重遮挡和开放集场景下难以同时生成高质量几何结构与准确姿态。为此,我们首先将去遮挡模型从3D物体生成中解耦,并利用图像数据集与自收集的去遮挡数据集,增强对多样化开放集遮挡模式的建模能力。其次,提出统一的姿态估计模型,融合自注意力与交叉注意力的全局与局部机制以提升精度。此外,构建了一个开放集3D场景数据集,进一步提升姿态估计模型的泛化能力。大量实验表明,该解耦框架在室内与开放集场景下均表现更优。代码与数据集已公开于https://idea-research.github.io/SceneMaker/。

原文摘要 · Abstract (English)

We propose a decoupled 3D scene generation framework called SceneMaker in this work. Due to the lack of sufficient open-set de-occlusion and pose estimation priors, existing methods struggle to simultaneously produce high-quality geometry and accurate poses under severe occlusion and open-set settings. To address these issues, we first decouple the de-occlusion model from 3D object generation, and enhance it by leveraging image datasets and collected de-occlusion datasets for much more diverse open-set occlusion patterns. Then, we propose a unified pose estimation model that integrates global and local mechanisms for both self-attention and cross-attention to improve accuracy. Besides, we construct an open-set 3D scene dataset to further extend the generalization of the pose estimation model. Comprehensive experiments demonstrate the superiority of our decoupled framework on both indoor and open-set scenes. Our codes and datasets is released at https://idea-research.github.io/SceneMaker/.

3D生成去遮挡姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。