首个多视角动态场景实时重建模型,支持单目视频生成子弹时间效果。
Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos
- 基于3D高斯点云,聚合所有上下文帧信息重建目标时刻场景。
- 单帧重建耗时150ms,静态与动态场景均达顶尖性能。
- 适用于影视特效、虚拟拍摄等需实时动态场景重建的场景。
近期静态前馈场景重建技术在高质量新视角合成方面取得显著进展,但这类模型在跨环境泛化性及动态内容处理上表现不佳。本文提出BTimer(BulletTimer),首个面向动态场景的运动感知前馈模型,实现动态场景的实时重建与新视角合成。该方法通过聚合所有上下文帧信息,在给定目标('子弹')时间戳下,以3D高斯点云形式重建完整场景。此设计使BTimer可同时利用静态与动态数据集,获得更强的可扩展性与泛化能力。面对任意单目动态视频,BTimer可在150毫秒内完成子弹时间场景重建,并在静态与动态场景数据集上达到领先水平,甚至优于基于优化的方法。
原文摘要 · Abstract (English)
Recent advancements in static feed-forward scene reconstruction have demonstrated significant progress in high-quality novel view synthesis. However, these models often struggle with generalizability across diverse environments and fail to effectively handle dynamic content. We present BTimer (short for BulletTimer), the first motion-aware feed-forward model for real-time reconstruction and novel view synthesis of dynamic scenes. Our approach reconstructs the full scene in a 3D Gaussian Splatting representation at a given target ('bullet') timestamp by aggregating information from all the context frames. Such a formulation allows BTimer to gain scalability and generalization by leveraging both static and dynamic scene datasets. Given a casual monocular dynamic video, BTimer reconstructs a bullet-time scene within 150ms while reaching state-of-the-art performance on both static and dynamic scene datasets, even compared with optimization-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。