arXiv:2411.19492cs.CVcs.LG2024-11ICCV被引 7

无需训练即可从单张图重建3D室内场景,支持未知物体和真实图像。

Diorama: Unleashing Zero-shot Single-view 3D Indoor Scene Modeling

  • 分解任务分别解决结构重建、形状检索、姿态估计与布局优化
  • 在合成与真实数据上均显著优于已有方法,支持跨域泛化
  • 适合需要快速生成3D场景的设计师与应用开发者

从RGB图像中利用CAD对象重建结构化3D场景,可实现高效紧凑的场景表示,同时保持组合性与交互性。现有方法依赖昂贵且不准确的真实标注,或可控但单调的合成数据,难以泛化到未见物体或领域。我们提出Diorama,首个零样本开放世界系统,仅需单视角RGB图像即可完成3D场景建模,无需端到端训练或人工标注。通过将问题分解为架构重建、3D形状检索、物体姿态估计与场景布局优化四个子任务,提出鲁棒且通用的解决方案。我们在合成与真实数据上评估系统性能,结果表明显著优于先前方法。此外,还验证了其对互联网图像及文本到场景任务的泛化能力。

原文摘要 · Abstract (English)

Reconstructing structured 3D scenes from RGB images using CAD objects unlocks efficient and compact scene representations that maintain compositionality and interactability. Existing works propose training-heavy methods relying on either expensive yet inaccurate real-world annotations or controllable yet monotonous synthetic data that do not generalize well to unseen objects or domains. We present Diorama, the first zero-shot open-world system that holistically models 3D scenes from single-view RGB observations without requiring end-to-end training or human annotations. We show the feasibility of our approach by decomposing the problem into subtasks and introduce robust, generalizable solutions to each: architecture reconstruction, 3D shape retrieval, object pose estimation, and scene layout optimization. We evaluate our system on both synthetic and real-world data to show we significantly outperform baselines from prior work. We also demonstrate generalization to internet images and the text-to-scene task.

3D重建零样本室内场景单视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。