arXiv:2609.06099cs.CV2026-09

用全景对齐技术让单目视频重建超出视野范围的4D场景

PASTEL: Panoramic Alignment for Monocular 4D Scene Reconstruction

论文配图:PASTEL: Panoramic Alignment for Monocular 4D Scene Reconstruction
图 1 · 摘自论文原文
  • 将不可见区域探索转为2D方向规划,降低视角搜索复杂度
  • 在DyCheck IPhone数据集上提升0.9dB的全图PSNR表现
  • 适合虚拟现实与具身智能中需要完整场景重建的场景

从随意拍摄的单目视频重建4D场景对虚拟现实和具身智能至关重要。现有方法通常无法恢复摄像头可见范围之外的区域。为此,我们提出一种新范式,结合单目输入的可见区域重建与超出可观测边界的内容生成。PASTEL通过全景对齐,将难以处理的3D不可见区域探索转化为可解的2D方向轨迹规划,将6-DoF视角搜索简化为带显式可见边界约束的2D方向搜索。在该全景空间中,方法能有效规划最大化拓展视野的相机路径,同时最小化视角偏差。实验表明,PASTEL不仅能合理外推输入视频可观测边界之外的场景内容,还显著提升单目4D重建性能,在DyCheck IPhone数据集上相比先前最先进方法提升0.9dB全图PSNR。

原文摘要 · Abstract (English)

Reconstructing 4D scenes from casually captured monocular video is vital for applications in virtual reality (VR) and embodied AI. Recent advances in 4D reconstruction and novel view synthesis have substantially propelled this capability. However, existing reconstruction methods generally cannot recover regions beyond visible camera limits. Consequently, we introduce a new paradigm that achieves 4D scene synthesis by combining visible-region reconstruction from monocular input with invisible-region generation beyond observable camera boundaries. We present Panoramic Alignment for Strategic Exploitation of Generative Priors (PASTEL). Specifically, PASTEL proposes panoramic scene alignment, a novel representation that reformulates the intractable 3D "invisible region" exploration into a tractable 2D directional trajectory planning. This is achieved by reducing the viewpoint planning from 6-DoF search to a 2D directional search with explicit visibility boundaries. By operating within this panoramic space, our method strategically identifies camera trajectories that maximize exploration beyond observable boundaries while minimizing viewpoint deviation. Experimental results show that PASTEL can not only extrapolate plausible scene content beyond the observable boundaries of input monocular videos, but also substantially boost monocular 4D reconstruction performance. PASTEL outperforms the previous state-of-the-art method by 0.9dB in full-image PSNR on the DyCheck IPhone dataset.

4D重建单目视觉全景对齐生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。