从极端运动模糊图像中重建3D场景,无需清晰参考图。
PRISM3D: Probabilistic Refinement and Robust Initialization for Physically Consistent Scene Modeling under Extreme Motion Blur

- 用深度稠密跟踪实现鲁棒初始化,突破传统SfM失效瓶颈。
- 结合概率密度与物理成像模型,解决初始稀疏噪声导致的优化发散。
- 支持多模态输入(RGB+事件数据),适合高动态复杂场景重建。
我们解决从极端运动模糊图像中进行盲式3D场景重建的逆问题,传统SfM方法在此类场景下失效。现有方法通常依赖不切实际的清晰图像监督。本文提出PRISM3D统一框架,可直接从严重退化输入中实现鲁棒重建。为克服缺乏可靠初始点的问题,提出基于VGGSfM的鲁棒初始化策略,恢复特征匹配失败时的全局拓扑结构。据我们所知,这是首次有效利用该范式从极端运动模糊中启动3D高斯点云。然而,该初始化结果稀疏且含噪,导致确定性优化发散。为此,我们提出耦合解法:采用马尔可夫链蒙特卡洛(MCMC)进行概率化几何稠密化,同时通过连续贝塞尔轨迹建模物理图像形成过程。此外,尽管PRISM3D已建立高度鲁棒的独立流程,但互补事件流的可用性为提升重建精度提供机会。为此,引入多模态扩展PRISM3D-E,无缝融合高时间分辨率事件数据作为结构先验以最大化几何恢复。由于现有数据集缺乏此类严重退化下的配对事件流,我们同步构建了PRISM3D-E基准数据集以促进严谨评估。大量实验表明,无论是单模态RGB框架还是其多模态扩展,均达到新最优性能。
原文摘要 · Abstract (English)
We address the inverse problem of blind 3D scene reconstruction from extremely motion-blurred images, a scenario where traditional Structure-from-Motion (SfM) pipelines fail. Existing approaches typically circumvent this bottleneck by relying on impractical sharp-image supervision. In this work, we introduce PRISM3D, a unified framework enabling robust reconstruction directly from severely degraded inputs. To overcome the lack of a reliable starting point, we propose a Robust Initialization strategy utilizing deep dense tracking method (VGGSfM) to recover global topology where feature matching fails. To the best of our knowledge, we are the first to effectively leverage this paradigm to bootstrap 3D Gaussian Splatting from extreme motion blur. However, while robust, this initialization yields sparse and noisy geometry that causes deterministic optimization to diverge. To resolve this, we propose a coupled solution driven by probability and physics: we adopt a probabilistic formulation for geometric densification via Markov Chain Monte Carlo (MCMC) to robustly populate the sparse priors, while simultaneously modeling physical image formation via continuous Bezier Trajectories. Furthermore, while PRISM3D establishes a highly robust standalone pipeline, the availability of complementary event streams offers an opportunity to push the reconstruction fidelity further. To exploit this, we introduce PRISM3D-E, a multi-modal (RGB + Events) extension that seamlessly integrates high-temporal-resolution events as structural priors to maximize geometric recovery. Because existing datasets lack paired event streams under such severe degradation, we concurrently contribute the PRISM3D-E Benchmark to facilitate rigorous evaluation. Extensive experiments demonstrate that both our standalone RGB framework and its multi-modal extension establish new state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。