仅用单目图像重建动态驾驶场景,无需位姿信息
FreeDriveRF: Monocular RGB Dynamic NeRF without Poses for Autonomous Driving via Point-Level Dynamic-Static Decoupling
- 在采样阶段通过语义监督分离动态与静态物体
- 利用光流约束动态建模,减少模糊和伪影
- 适合自动驾驶中无位姿依赖的实时场景理解
自动驾驶中的动态场景重建可更精准感知复杂环境变化。动态神经辐射场(NeRF)近期展现出强大建模能力,但多数方法依赖高精度位姿输入和多传感器数据,增加系统复杂性。为此,我们提出FreeDriveRF,仅使用连续RGB图像即可重建动态驾驶场景,无需位姿输入。创新性地在早期采样阶段基于语义监督实现动态-静态解耦,缓解图像模糊与伪影问题。针对单目相机下的运动与遮挡挑战,引入光流引导的动态物体渲染一致性损失,更好约束动态建模过程。同时结合估计的动态光流约束位姿优化,提升无界场景重建的稳定性与精度。在KITTI和Waymo数据集上的大量实验表明,本方法在自动驾驶动态场景建模上表现优异。
原文摘要 · Abstract (English)
Dynamic scene reconstruction for autonomous driving enables vehicles to perceive and interpret complex scene changes more precisely. Dynamic Neural Radiance Fields (NeRFs) have recently shown promising capability in scene modeling. However, many existing methods rely heavily on accurate poses inputs and multi-sensor data, leading to increased system complexity. To address this, we propose FreeDriveRF, which reconstructs dynamic driving scenes using only sequential RGB images without requiring poses inputs. We innovatively decouple dynamic and static parts at the early sampling level using semantic supervision, mitigating image blurring and artifacts. To overcome the challenges posed by object motion and occlusion in monocular camera, we introduce a warped ray-guided dynamic object rendering consistency loss, utilizing optical flow to better constrain the dynamic modeling process. Additionally, we incorporate estimated dynamic flow to constrain the pose optimization process, improving the stability and accuracy of unbounded scene reconstruction. Extensive experiments conducted on the KITTI and Waymo datasets demonstrate the superior performance of our method in dynamic scene modeling for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。