arXiv:2411.19141cs.CVcs.AI2024-11ICCV被引 6

用Transformer融合视觉与运动信息,提升单目视频动目标分割效果

On Moving Object Segmentation from Monocular Video with Transformers

  • 设计M3Former架构,结合视觉与运动特征进行分割
  • 验证2D/3D运动表示对分割性能的关键影响
  • 强调多样数据集训练对达到顶尖性能至关重要

从单个移动摄像头的视频中进行运动物体检测与分割是一项挑战性任务,需理解识别、运动和三维几何关系。将识别与重建结合本质上是多模态融合问题,需融合外观与运动特征以实现分类与分割。本文提出一种新型单目运动分割融合架构M3Former,利用Transformer在分割与多模态融合中的强大表现。由于从单目视频重建运动属于病态问题,我们系统分析了不同2D与3D运动表示对该任务的影响及其对分割性能的重要性。最后,分析了训练数据的影响,表明在Kitti与Davis数据集上取得最先进性能需依赖多样化数据集。

原文摘要 · Abstract (English)

Moving object detection and segmentation from a single moving camera is a challenging task, requiring an understanding of recognition, motion and 3D geometry. Combining both recognition and reconstruction boils down to a fusion problem, where appearance and motion features need to be combined for classification and segmentation. In this paper, we present a novel fusion architecture for monocular motion segmentation - M3Former, which leverages the strong performance of transformers for segmentation and multi-modal fusion. As reconstructing motion from monocular video is ill-posed, we systematically analyze different 2D and 3D motion representations for this problem and their importance for segmentation performance. Finally, we analyze the effect of training data and show that diverse datasets are required to achieve SotA performance on Kitti and Davis.

运动分割Transformer单目视频多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。