用前馈网络高效重建动态3D场景,无需复杂优化。
MoRe: Motion-aware Feed-forward 4D Reconstruction Transformer
- 通过注意力强制分离运动与静态结构,实现端到端重建。
- 在多个基准上达到高质量动态重建,速度远超传统方法。
- 适合实时动态场景重建应用,如AR/VR和自动驾驶。
从单目视频中重建动态4D场景仍具挑战性,因运动物体会干扰相机位姿估计。现有优化方法虽能缓解此问题,但大多计算开销大,难以实现实时应用。为此,我们提出MoRe,一种前馈式4D重建网络,可高效恢复动态3D场景。基于强大的静态重建主干网络,MoRe采用注意力强制策略,将动态运动与静态结构解耦。为增强鲁棒性,模型在包含动态与静态场景的大规模多样化数据集上进行微调。此外,分组因果注意力机制捕捉时间依赖性,并适应各帧间不同长度的标记,确保几何重建的时间一致性。在多个基准上的大量实验表明,MoRe在保持高重建质量的同时,展现出卓越的效率。
原文摘要 · Abstract (English)
Reconstructing dynamic 4D scenes remains challenging due to the presence of moving objects that corrupt camera pose estimation. Existing optimization methods alleviate this issue with additional supervision, but they are mostly computationally expensive and impractical in real-time applications. To address these limitations, we propose MoRe, a feedforward 4D reconstruction network that efficiently recovers dynamic 3D scenes from monocular videos. Built upon a strong static reconstruction backbone, MoRe employs an attention-forcing strategy to disentangle dynamic motion from static structure. To further enhance robustness, we fine-tune the model on large-scale, diverse datasets encompassing both dynamic and static scenes. Moreover, our grouped causal attention captures temporal dependencies and adapts to varying token lengths across frames, ensuring temporally coherent geometry reconstruction. Extensive experiments on multiple benchmarks demonstrate that MoRe achieves high-quality dynamic reconstructions with exceptional efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。