arXiv:2608.23930cs.CV2026-08

从单张图生成完整3D场景,让物体自然融入统一空间框架。

SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image

论文配图:SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image
图 1 · 摘自论文原文
  • 用可学习查询引导预训练模型生成带姿态的完整物体网格
  • 在3D-FUTURE上多项指标领先,物体与场景重建均表现优异
  • 适合自动驾驶和具身智能等需要真实场景理解的领域

单图像3D场景重建需补全部分可见物体,并将其合理放置于统一观测对齐的场景坐标系中。现有基于物体生成先验的方法虽能有效补全,但其输出通常以物体为中心、尺度归一化,与场景重建所需的统一坐标系存在根本性表征差距。本文提出SceneReGen,将场景重建重构为在共享观测对齐场景帧中生成并组装完整物体资产的过程。通过选择性姿态因子分解:直接在生成网格中编码物体观测方向,而平移与尺度则由实例级和全局场景证据估计。给定场景图像和实例掩码后,几何编码器提取密集特征;可学习形状查询引导基于DiT的3D生成器输出带观测姿态的完整网格,位置查询融合物体与场景特征,实现精准组装。在3D-FUTURE评估子集上,SceneReGen在场景级CD、F-Score及3D边界框IoU上均最优,物体级CD并列第一,物体级F-Score排名第二。自动驾驶与具身智能场景的定性结果进一步展示了以资产为中心重建的潜力。

原文摘要 · Abstract (English)

Single-image 3D scene reconstruction must complete partially observed objects and place them coherently in a shared observation-aligned scene frame. Object-level generative priors offer strong completion ability, but their centered, scale-normalized outputs are typically expressed in an object frame, creating a fundamental representation gap between object generation and scene reconstruction. We introduce SceneReGen, a generative reconstruction framework that reinterprets scene reconstruction as the generation and assembly of complete object assets in a shared observation-aligned scene frame. SceneReGen addresses the generation-reconstruction gap through selective pose factorization: each object's observed orientation is encoded directly in the generated mesh, while translation and scale are estimated from instance-level and global scene evidence. Given a scene image and instance masks, a geometry encoder extracts dense cues; learnable shape queries condition a pretrained DiT-based 3D generator to produce complete meshes in their observed orientations, while position queries fuse object and scene features to assemble them in the shared frame. On the 3D-FUTURE evaluation subset, SceneReGen achieves the best scene-level CD, scene-level F-Score, and 3D bounding-box IoU among the evaluated methods, ties the best object-level CD, and ranks second in object-level F-Score. Qualitative outputs in autonomous-driving and embodied-AI scenes further illustrate the potential of asset-centric reconstruction beyond indoor furniture.

3D生成单图重建场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。