通过跨模态令牌调制,提升视频物体分割中外观与运动信息的融合效果。
CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation

- 设计跨模态令牌调制机制,增强外观与运动特征间的交互。
- 在所有公开基准上达到当前最优性能,无需额外标注数据。
- 采用令牌掩码策略,提升训练效率,避免单纯堆叠模型复杂度。
最近的无监督视频对象分割研究显示,融合外观和运动线索的双流架构具有巨大潜力。然而,充分挖掘这两种互补信息需有效建模其相互依赖关系。本文提出跨模态令牌调制,一种新型方法,旨在加强外观与运动线索间的互动。该方法在各模态的令牌间建立密集连接,通过关系变换器块实现高效的模内与跨模态信息传播。为提升学习效率,引入令牌掩码策略,克服仅靠增加模型复杂度的局限性。本方法在所有公开基准上均取得当前最佳表现。
原文摘要 · Abstract (English)
Recent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully leveraging these complementary sources of information requires effectively modeling their interdependencies. In this paper, we introduce cross-modality token modulation, a novel approach designed to strengthen the interaction between appearance and motion cues. Our method establishes dense connections between tokens from each modality, enabling efficient intra-modal and inter-modal information propagation through relation transformer blocks. To improve learning efficiency, we incorporate a token masking strategy that addresses the limitations of relying solely on increased model complexity. Our approach achieves state-of-the-art performance across all public benchmarks, outperforming existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。