用音频引导视觉,防止音视频增量学习遗忘旧知识
Listen, Look, and Learn: Learning Without Forgetting through SAM-Audio

- 音频特征动态引导视觉表示,实现跨模态协同学习
- 在多个音视频增量学习基准上超越现有方法,最高提升6.2%
- 适合研究多模态持续学习的开发者和研究人员
类别增量学习(CIL)旨在不遗忘已有知识的前提下持续学习新类别。尽管近年来多模态CIL取得进展,但音视频场景仍研究不足。我们发现,像SAM-Audio这样的基础多模态模型虽具备丰富的静态先验,但在增量学习中表现不佳。本文将SAM-Audio的密集音视频表征融入CIL框架,提出一种新型引导注意力机制,使音频特征上下文化地指导视觉表示。为缓解灾难性遗忘,引入特征级与输出级双层知识蒸馏。在多个音视频CIL基准上的实验表明,该方法显著优于当前最优方法。
原文摘要 · Abstract (English)
Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. While recent CIL advances have spurred significant interest across various modalities, the audio-visual setting remains underexplored. Furthermore, although foundational multimodal models like SAM-Audio encapsulate rich static priors, our empirical analysis reveals that these representations struggle in incremental settings. This work bridges this gap by integrating SAM-Audio's audio-visual priors into the CIL setting. Specifically, we leverage its dense audio and visual representations and employ a novel guided attention strategy where the audio features contextually guide the visual representations. To further mitigate catastrophic forgetting, we introduce dual-level distillation objectives at both the feature and logit levels. Extensive evaluations on audio-visual CIL benchmarks demonstrate that our approach consistently outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。