arXiv:2504.05904cs.CVcs.LG2025-04

通过视觉显著性与运动引导,提升无监督视频目标分割精度

Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation

  • 设计双分支结构,分离共性与运动特异性特征
  • 在三个数据集上达到最高89.2%的J&F得分
  • 无需额外输入,利用模型自身显著性信息优化分割

当前主流无监督视频目标分割方法多采用双编码器或单编码器结构,但难以平衡运动与外观特征的关系。即使使用复杂融合模块,仍因特征提取不优而影响整体性能。此外,光流质量受场景影响大,仅依赖光流难以实现高质量分割。为此,本文提出Saliency-Motion引导的Trunk-Collateral网络(SMTC-Net),更好地平衡运动与外观关系,并引入模型内在显著性信息以增强分割效果。具体地,共享主干网络捕捉运动与外观共性,旁路分支学习运动特异性;同时设计内在显著性引导精炼模块(ISRM),高效利用模型自身显著性信息,对高层特征进行精炼并提供像素级融合指导,无需额外输入。实验表明,SMTC-Net在三个标准UVOS数据集(DAVIS-16: 89.2% J&F,YouTube-Objects: 76% J,FBMS: 86.4% J)和四个视频显著性检测基准上均达领先水平,验证了其有效性与优越性。

原文摘要 · Abstract (English)

Recent mainstream unsupervised video object segmentation (UVOS) motion-appearance approaches use either the bi-encoder structure to separately encode motion and appearance features, or the uni-encoder structure for joint encoding. However, these methods fail to properly balance the motion-appearance relationship. Consequently, even with complex fusion modules for motion-appearance integration, the extracted suboptimal features degrade the models' overall performance. Moreover, the quality of optical flow varies across scenarios, making it insufficient to rely solely on optical flow to achieve high-quality segmentation results. To address these challenges, we propose the Saliency-Motion guided Trunk-Collateral Network (SMTC-Net), which better balances the motion-appearance relationship and incorporates model's intrinsic saliency information to enhance segmentation performance. Specifically, considering that optical flow maps are derived from RGB images, they share both commonalities and differences. Accordingly, we propose a novel Trunk-Collateral structure for motion-appearance UVOS. The shared trunk backbone captures the motion-appearance commonality, while the collateral branch learns the uniqueness of motion features. Furthermore, an Intrinsic Saliency guided Refinement Module (ISRM) is devised to efficiently leverage the model's intrinsic saliency information to refine high-level features, and provide pixel-level guidance for motion-appearance fusion, thereby enhancing performance without additional input. Experimental results show that SMTC-Net achieved state-of-the-art performance on three UVOS datasets ( 89.2% J&F on DAVIS-16, 76% J on YouTube-Objects, 86.4% J on FBMS ) and four standard video salient object detection (VSOD) benchmarks with the notable increase, demonstrating its effectiveness and superiority over previous methods.

视频分割无监督学习显著性引导运动感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。