arXiv:2606.12826cs.CVcs.AI2026-06

解耦视觉与事件特征,提升小目标动态分割精度

DIMOS: Disentangling Instance-level Moving Object Segmentation

论文配图:DIMOS: Disentangling Instance-level Moving Object Segmentation
图 1 · 摘自论文原文
  • 分离图像与事件模态中的外观和运动信息,增强特征密度
  • 多粒度跨模态对齐实现时空特征有效融合,小目标分割提升明显
  • 适合低光、高速运动等复杂场景下的动态目标分割任务

移动实例分割(MIS)因其在交通监控、自动驾驶和动物追踪中的广泛应用而受到越来越多关注。事件相机记录异步亮度变化,具备高时间分辨率和动态范围,对运动信息极为敏感。通过融合事件与图像特征,事件中的运动线索可补充图像的空间细节,从而提升MIS性能。然而,现有多模态MIS方法在分割小尺寸移动实例时仍存在困难,因为事件相机在有限分辨率下常产生稀疏特征。此外,事件特征中外观属性与运动线索相互纠缠,进一步限制了跨模态融合效果。为此,本文提出一种双解耦特征提取框架,分别在图像与事件模态内分离并提取外观与运动信息,提升特征密度;随后引入多粒度跨模态对齐机制,对齐模态间分布与语义一致的特征,实现更有效的融合,获得丰富时空细节。实验结果表明,本方法在多模态MIS任务中达到当前最优性能,尤其在快速运动和低光照等挑战性条件下对小目标的分割表现显著提升。

原文摘要 · Abstract (English)

Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking. Event cameras record asynchronous brightness changes, providing high temporal resolution and dynamic range, which makes them highly sensitive to motion information. By fusing event and image features, motion cues from events can complement spatial details from images, enhancing the performance of MIS. However, current multimodal MIS methods still struggle to segment small moving instances, as event cameras often yield sparse features under limited resolution. Moreover, event features entangle appearance attributes with motion cues, which further restricts effective cross-modal fusion. To address these challenges, we first propose a dual-disentangling feature extraction framework that separates and extracts appearance and motion information within both image and event modalities, thereby improving feature density. Subsequently, a multi-granularity cross-modal alignment is introduced to align distributionally and semantically consistent features across modalities, enabling more effective fusion with rich spatial and temporal details. The experiment results demonstrate that our method achieves state-of-the-art performance in multimodal MIS, especially for small instances under challenging conditions such as fast motion and low-light settings.

动态分割事件相机多模态融合小目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。