arXiv:2502.12975cs.CV2025-02IJCV被引 12

融合图像与事件数据,实现单图下实例级动态物体分割

Instance-Level Moving Object Segmentation from a Single Image with Events

  • 用图像纹理与事件运动信息互补建模
  • 在多个数据集上达到领先性能,提升显著
  • 适合做事件感知视觉系统的研究者参考

动态场景中的移动物体分割对理解复杂运动至关重要,但难点在于同时捕捉空间纹理结构和时间运动线索。基于视频帧的方法因难以区分相机运动与物体运动而受限。近年来,事件相机因其对运动的敏感性被引入,但其缺乏密集纹理信息导致像素级掩码分割困难。为克服单模态局限,我们提出首个整合图像与事件数据的实例级移动物体分割框架。模型通过隐式跨模态掩码注意力增强、显式对比特征学习和光流引导运动增强,分别利用图像中的稠密纹理和事件中的丰富运动信息。通过增强后的特征,将掩码分割与运动分类解耦,以处理任意数量独立运动物体。在多个数据集上的广泛评估、不同输入设置的消融实验及实时效率分析表明,该方法首次实现了图像与事件数据融合的实用部署,为未来事件驱动的运动相关研究提供新思路。代码与预训练权重已开源。

原文摘要 · Abstract (English)

Moving object segmentation plays a crucial role in understanding dynamic scenes involving multiple moving objects, while the difficulties lie in taking into account both spatial texture structures and temporal motion cues. Existing methods based on video frames encounter difficulties in distinguishing whether pixel displacements of an object are caused by camera motion or object motion due to the complexities of accurate image-based motion modeling. Recent advances exploit the motion sensitivity of novel event cameras to counter conventional images' inadequate motion modeling capabilities, but instead lead to challenges in segmenting pixel-level object masks due to the lack of dense texture structures in events. To address these two limitations imposed by unimodal settings, we propose the first instance-level moving object segmentation framework that integrates complementary texture and motion cues. Our model incorporates implicit cross-modal masked attention augmentation, explicit contrastive feature learning, and flow-guided motion enhancement to exploit dense texture information from a single image and rich motion information from events, respectively. By leveraging the augmented texture and motion features, we separate mask segmentation from motion classification to handle varying numbers of independently moving objects. Through extensive evaluations on multiple datasets, as well as ablation experiments with different input settings and real-time efficiency analysis of the proposed framework, we believe that our first attempt to incorporate image and event data for practical deployment can provide new insights for future work in event-based motion related works. The source code with model training and pre-trained weights is released at https://npucvr.github.io/EvInsMOS

实例分割事件相机多模态融合动态场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。