arXiv:2603.08967cs.CVeess.AS2026-03

首个免样本的音视频分割持续学习基准,解决动态环境下的长期感知挑战。

Can You Hear, Localize, and Segment Continually? An Exemplar-Free Continual Learning Benchmark for Audio-Visual Segmentation

  • 提出无样本持续学习框架,支持单源与多源数据流
  • 基线模型ATLAS通过音频引导特征调制提升跨模态融合效果
  • 低秩锚定机制有效缓解遗忘,适合长期部署场景

音视频分割(AVS)旨在通过联合学习音频与视觉信号,生成视频中发声物体的像素级掩码。然而真实环境具有动态性,导致音视频分布随时间演变,现有AVS系统因假设静态训练设置而面临挑战。为此,本文首次提出免样本的持续学习基准,涵盖单源与多源AVS数据集上的四种学习协议。进一步提出强基线模型ATLAS,利用音频引导的预融合条件调节,在跨模态注意力前通过投影音频上下文调制视觉特征通道。为缓解灾难性遗忘,引入低秩锚定(LRA)机制,基于损失敏感度稳定适应后的权重。大量实验表明,该方法在多种持续学习场景下表现优异,为长期音视频感知奠定基础。代码已公开。

原文摘要 · Abstract (English)

Audio-Visual Segmentation (AVS) aims to produce pixel-level masks of sound producing objects in videos, by jointly learning from audio and visual signals. However, real-world environments are inherently dynamic, causing audio and visual distributions to evolve over time, which challenge existing AVS systems that assume static training settings. To address this gap, we introduce the first exemplar-free continual learning benchmark for Audio-Visual Segmentation, comprising four learning protocols across single-source and multi-source AVS datasets. We further propose a strong baseline, ATLAS, which uses audio-guided pre-fusion conditioning to modulate visual feature channels via projected audio context before cross-modal attention. Finally, we mitigate catastrophic forgetting by introducing Low-Rank Anchoring (LRA), which stabilizes adapted weights based on loss sensitivity. Extensive experiments demonstrate competitive performance across diverse continual scenarios, establishing a foundation for lifelong audio-visual perception. Code is available at${}^{*}$\footnote{Paper under review} - \hyperlink{https://gitlab.com/viper-purdue/atlas}{https://gitlab.com/viper-purdue/atlas} \keywords{Continual Learning \and Audio-Visual Segmentation \and Multi-Modal Learning}

音视频分割持续学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。