arXiv:2507.21588cs.AIcs.CV2025-07ICCV被引 3

提出分阶段提示调优方法,让音视频多任务模型持续学习不遗忘。

Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning

  • 三阶段设计:共享-动态-独立提示,逐步增强跨任务理解
  • 在四个音视频任务上按不同顺序增量学习均达最优性能
  • 适合需要长期更新的音视频多任务系统开发

音视频多任务增量学习旨在无需联合训练所有任务的情况下持续学习多个音视频任务。核心挑战在于如何在保留旧任务知识的同时有效学习新任务。为此,我们提出三阶段渐进式稳态与可塑提示(PHP)方法。浅层阶段设计任务共享模态聚合适配器,促进跨任务、跨模态的音视频表示学习,增强任务间共享理解;中层阶段提出任务特定模态共享动态生成适配器,构建针对各任务但跨模态通用的提示,平衡知识保持与多任务迁移能力;深层阶段引入任务特定模态无关提示,进一步细化每个任务和模态的表征能力。通过三阶段协同,PHP在保留任务专属提示的同时适应共享参数以应对新任务,有效平衡知识共享与专属性。该方法在四种任务(AVE、AVVP、AVS、AVQA)的不同顺序下均达到当前最优表现。

原文摘要 · Abstract (English)

Audio-visual multi-task incremental learning aims to continuously learn from multiple audio-visual tasks without the need for joint training on all tasks. The challenge of the problem is how to preserve the old task knowledge while facilitating the learning of new task with previous experiences. To address these challenges, we introduce a three-stage Progressive Homeostatic and Plastic audio-visual prompt (PHP) method. In the shallow phase, we design the task-shared modality aggregating adapter to foster cross-task and cross-modal audio-visual representation learning to enhance shared understanding between tasks. In the middle phase, we propose the task-specific modality-shared dynamic generating adapter, which constructs prompts that are tailored to individual tasks while remaining general across modalities, which balances the models ability to retain knowledge against forgetting with its potential for versatile multi-task transferability. In the deep phase, we introduce the task-specific modality-independent prompts to further refine the understand ability by targeting individual information for each task and modality. By incorporating these three phases, PHP retains task-specific prompts while adapting shared parameters for new tasks to effectively balance knowledge sharing and specificity. Our method achieves SOTA performance in different orders of four tasks (AVE, AVVP, AVS and AVQA). Our code can be available at https://github.com/ENJOY-Yin-jiong/PHP.

音视频增量学习提示调优多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。