arXiv:2604.10578cs.CV2026-04被引 2

用视频扩散模型修复360度室内场景,实现全局一致的高保真重建。

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models

论文配图:Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models
图 1 · 摘自论文原文
  • 结合3D高斯泼溅与全景视频扩散模型,分步恢复并优化场景。
  • 在15,000对全景视频数据上训练,显著提升远距离相机探索能力。
  • 适合做虚拟现实与具身智能的高质量3D场景生成研究者。

日益增长的具身智能与虚拟现实应用需求,推动了从稀疏输入生成高质量3D室内场景的研究。然而,现有方法难以在大范围未知区域推断大量缺失几何结构,且常导致全局不一致,仅生成局部合理但整体失真的结果。本文提出Rein3D框架,通过将显式3D高斯泼溅(3DGS)与视频扩散模型提供的时序一致性先验相结合,重建完整的360度室内环境。该方法采用“还原-精炼”范式:利用径向探索策略,从初始点沿轨迹渲染不完美的全景视频,有效揭示被遮挡区域;再通过全景视频到视频扩散模型还原序列,并经视频超分辨率增强,生成高保真几何与纹理;最终,这些精炼后的视频作为伪真值,用于更新全局3D高斯场。为支持此任务,我们构建了包含超过15,000对干净与退化全景视频的PanoV2V-15K数据集。实验表明,Rein3D可生成逼真且全局一致的3D场景,在长距离相机探索方面显著优于现有基线。

原文摘要 · Abstract (English)

The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. However, existing approaches struggle to infer massive amounts of missing geometry in large unseen areas while maintaining global consistency, often producing locally plausible but globally inconsistent reconstructions. We present Rein3D, a framework that reconstructs full 360-degree indoor environments by coupling explicit 3D Gaussian Splatting (3DGS) with temporally coherent priors from video diffusion models. Our approach follows a "restore-and-refine" paradigm: we employ a radial exploration strategy to render imperfect panoramic videos along trajectories starting from the origin, effectively uncovering occluded regions from a coarse 3DGS initialization. These sequences are restored by a panoramic video-to-video diffusion model and further enhanced via video super-resolution to synthesize high-fidelity geometry and textures. Finally, these refined videos serve as pseudo-ground truths to update the global 3D Gaussian field. To support this task, we construct PanoV2V-15K, a dataset of over 15K paired clean and degraded panoramic videos for diffusion-based scene restoration. Experiments demonstrate that Rein3D produces photorealistic and globally consistent 3D scenes and significantly improves long-range camera exploration compared with existing baselines.

3D生成视频扩散具身智能高斯泼溅

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。