UFO统一前馈与优化方法,实现长序列高精度4D动态场景重建。
UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
- 引入递归框架,通过可见性筛选高效处理长序列输入。
- 16秒驾驶日志重建仅需0.5秒,视觉与几何精度更优。
- 适合自动驾驶仿真与闭环学习,尤其长时序场景建模。
动态驾驶场景重建对自动驾驶仿真和闭环学习至关重要。尽管近期前馈方法在3D重建上表现良好,但在长序列驾驶场景中因序列长度呈二次复杂度且难以建模长时间动态物体而受限。本文提出UFO,一种融合优化与前馈方法优势的新型递归范式,实现高效长时序4D重建。该方法维护一个可迭代更新的4D场景表示,随新观测持续优化,利用基于可见性的过滤机制选择关键场景标记,提升长序列处理效率。针对动态物体,提出物体位姿引导建模,支持精准长时运动捕捉。在Waymo Open Dataset上的实验表明,该方法在多种序列长度下显著优于单场景优化及现有前馈方法。特别地,可在0.5秒内完成16秒驾驶日志的重建,同时保持卓越的视觉质量与几何精度。
原文摘要 · Abstract (English)
Dynamic driving scene reconstruction is critical for autonomous driving simulation and closed-loop learning. While recent feed-forward methods have shown promise for 3D reconstruction, they struggle with long-range driving sequences due to quadratic complexity in sequence length and challenges in modeling dynamic objects over extended durations. We propose UFO, a novel recurrent paradigm that combines the benefits of optimization-based and feed-forward methods for efficient long-range 4D reconstruction. Our approach maintains a 4D scene representation that is iteratively refined as new observations arrive, using a visibility-based filtering mechanism to select informative scene tokens and enable efficient processing of long sequences. For dynamic objects, we introduce an object pose-guided modeling approach that supports accurate long-range motion capture. Experiments on the Waymo Open Dataset demonstrate that our method significantly outperforms both per-scene optimization and existing feed-forward methods across various sequence lengths. Notably, our approach can reconstruct 16-second driving logs within 0.5 second while maintaining superior visual quality and geometric accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。