arXiv:2411.18650cs.CV2024-11ICCV被引 20

提出RoMo方法,用运动分割提升动态场景的相机标定精度

RoMo: Robust Motion Segmentation Improves Structure from Motion

论文配图:RoMo: Robust Motion Segmentation Improves Structure from Motion
图 1 · 摘自论文原文
  • 结合光流、对极几何与预训练模型进行迭代运动分割
  • 在动态场景下相机标定误差显著降低,优于现有方法
  • 适合需要高精度相机参数的4D重建应用

从单目随意拍摄视频中重建和生成4D场景已取得显著进展,但这些任务高度依赖已知相机位姿。而基于结构从运动(SfM)求解位姿时,往往受限于难以可靠分离视频中的静态与动态部分。本文提出一种新型视频运动分割方法RoMo,用于识别相对于固定世界坐标系移动的场景成分。该方法通过迭代融合光学流、对极几何线索与预训练视频分割模型,显著优于无监督基线及在合成数据上训练的有监督基线。更重要的是,将现成SfM管线与我们的分割掩码结合,在含动态内容的场景相机标定任务上达到新基准,性能超越现有方法显著。

原文摘要 · Abstract (English)

There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion (SfM) often depends on robustly separating static from dynamic parts of a video. The lack of a robust solution to this problem limits the performance of SfM camera-calibration pipelines. We propose a novel approach to video-based motion segmentation to identify the components of a scene that are moving w.r.t. a fixed world frame. Our simple but effective iterative method, RoMo, combines optical flow and epipolar cues with a pre-trained video segmentation model. It outperforms unsupervised baselines for motion segmentation as well as supervised baselines trained from synthetic data. More importantly, the combination of an off-the-shelf SfM pipeline with our segmentation masks establishes a new state-of-the-art on camera calibration for scenes with dynamic content, outperforming existing methods by a substantial margin.

运动分割4D重建相机标定SfM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。