arXiv:2508.15769cs.CVcs.AI2025-08中稿 · 3DV 2026被引 44

单张图像一键生成带位置关系的3D场景,无需迭代优化

SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass

  • 输入图像与物体掩码,一次前向传播生成多物体3D几何与纹理
  • 融合视觉与几何特征,直接预测物体空间位置,速度更快
  • 支持多图输入提升效果,适合虚拟现实与机器人应用

3D内容生成因在虚拟现实、增强现实及具身智能中的关键应用而备受关注。本文提出SceneGen,一种新框架,仅需输入一张场景图像及其对应物体掩码,即可一次性生成多个带有几何与纹理的3D资产。其核心优势在于无需额外优化或资产检索。我们设计了一种新型特征聚合模块,整合视觉与几何编码器的局部与全局信息,并结合位置头,在单次前向传播中完成3D资产生成及其相对空间位置预测。该模型在多图像输入下仍表现更优,即使训练仅基于单图。大量定量与定性评估验证了方法的高效性与鲁棒性。我们相信这一范式为高质量3D内容生成提供了新思路,有望推动下游任务的实际应用。代码与模型将公开于:https://mengmouxu.github.io/SceneGen。

原文摘要 · Abstract (English)

3D content generation has recently attracted significant research interest, driven by its critical applications in VR/AR and embodied AI. In this work, we tackle the challenging task of synthesizing multiple 3D assets within a single scene image. Concretely, our contributions are fourfold: (i) we present SceneGen, a novel framework that takes a scene image and corresponding object masks as input, simultaneously producing multiple 3D assets with geometry and texture. Notably, SceneGen operates with no need for extra optimization or asset retrieval; (ii) we introduce a novel feature aggregation module that integrates local and global scene information from visual and geometric encoders within the feature extraction module. Coupled with a position head, this enables the generation of 3D assets and their relative spatial positions in a single feedforward pass; (iii) we demonstrate SceneGen's direct extensibility to multi-image input scenarios. Despite being trained solely on single-image inputs, our architecture yields improved generation performance when multiple images are provided; and (iv) extensive quantitative and qualitative evaluations confirm the efficiency and robustness of our approach. We believe this paradigm offers a novel solution for high-quality 3D content generation, potentially advancing its practical applications in downstream tasks. The code and model will be publicly available at: https://mengmouxu.github.io/SceneGen.

3D生成单图重建场景生成前向生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。