解决音频视觉分割中模态混淆问题,提升持续学习效果
Taming Modality Entanglement in Continual Audio-Visual Segmentation
- 通过多模态样本选择增强模态一致性,缓解语义漂移
- 设计碰撞机制增加易混淆类样本重放频率,减少共现混淆
- 适用于需要持续学习细粒度音视频分割的场景
近期多模态持续学习取得进展,但在细粒度任务中仍面临模态纠缠问题。本文提出新的持续音频-视觉分割(CAVS)任务,旨在通过音频引导连续分割新类别。分析发现两大挑战:1)多模态语义漂移,即发声物体在后续任务中被误标为背景;2)共现混淆,频繁共现类别易被混淆。为此,提出基于碰撞的多模态重放(CMR)框架:针对语义漂移,采用多模态样本选择(MSS)策略选取高模态一致性样本进行重放;针对共现混淆,设计碰撞式重放机制,在训练中动态提高易混淆类别的重放频率。构建三个音视频增量场景验证方法有效性。大量实验表明,本方法显著优于单模态持续学习方法。
原文摘要 · Abstract (English)
Recently, significant progress has been made in multi-modal continual learning, aiming to learn new tasks sequentially in multi-modal settings while preserving performance on previously learned ones. However, existing methods mainly focus on coarse-grained tasks, with limitations in addressing modality entanglement in fine-grained continual learning settings. To bridge this gap, we introduce a novel Continual Audio-Visual Segmentation (CAVS) task, aiming to continuously segment new classes guided by audio. Through comprehensive analysis, two critical challenges are identified: 1) multi-modal semantic drift, where a sounding objects is labeled as background in sequential tasks; 2) co-occurrence confusion, where frequent co-occurring classes tend to be confused. In this work, a Collision-based Multi-modal Rehearsal (CMR) framework is designed to address these challenges. Specifically, for multi-modal semantic drift, a Multi-modal Sample Selection (MSS) strategy is proposed to select samples with high modal consistency for rehearsal. Meanwhile, for co-occurence confusion, a Collision-based Sample Rehearsal (CSR) mechanism is designed, allowing for the increase of rehearsal sample frequency of those confusable classes during training process. Moreover, we construct three audio-visual incremental scenarios to verify effectiveness of our method. Comprehensive experiments demonstrate that our method significantly outperforms single-modal continual learning methods. Code can be seen at https://github.com/cqu-student/CAVS-CMR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。