arXiv:2603.04128cs.CVcs.AI2026-03

解决音视频多任务学习中的负迁移问题,实现高效统一建模。

Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

  • 通过显式协作机制,从数据与模型双角度缓解音视频任务异质性。
  • 在17个数据集、7项任务上实现近88%的任务性能超越单任务基线。
  • 适用于多模态理解场景,尤其适合需要统一建模的复杂视听应用。

构建统一的音视频大语言模型(AV-LLMs)对多模态智能至关重要。尽管指令微调可赋予预训练模型多任务能力,但传统多任务统一方法常因严重负迁移导致近55%的任务性能下降。我们发现其根源在于音视频任务异质性——任务粒度差异大、能力需求不一,联合训练时引发干扰。为此,提出Crab⁺模型,通过数据与模型双路径实现显式协作:数据侧引入包含约22.2万样本、覆盖17个数据集和7项任务的AV-UIE v2统一指令数据集,支持跨任务关系建模;模型侧设计统一接口对齐异构任务形式,并提出交互感知LoRA(I-LoRA),通过动态路由显式建模任务间关系,缓解参数干扰。实验表明,Crab⁺覆盖任务更广,且在多个基准上超越专用模型。成功逆转负迁移趋势,在近88%的任务中实现正迁移。结果在多种AV-LLM范式下稳定,经可视化验证,为实现整体音视频场景理解迈出坚实一步。

原文摘要 · Abstract (English)

Developing Audio-Visual Large Language Models (AV-LLMs) for unified scene understanding is pivotal in multimodal intelligence. While instruction tuning enables pre-trained models with multi-task abilities, we observe that conventional multi-task unification methods often suffer from severe negative transfer, where nearly 55% of tasks degrade compared to single-task training. We attribute this phenomenon to audio-visual task heterogeneity, characterized by disparate task granularity and divergent capability demands, which lead to negative interference under joint training. To tackle this, we present Crab$^{+}$, a scalable and unified audio-visual scene understanding model that addresses task heterogeneity through explicit cooperation from both data and model perspectives. On the data side, we introduce AV-UIE v2, a comprehensive Audio-Visual Unified Instruction-tuning dataset with Explicit reasoning processes. It contains approximately 222K samples spanning 17 datasets and 7 tasks, enabling the model to capture cross-task relationships at different levels of granularity. On the model side, we design a unified interface to align heterogeneous task formulations, and propose Interaction-aware LoRA (I-LoRA), which explicitly models inter-task relationships via dynamic routing to coordinate distinct audio-visual interaction patterns, mitigating parameter interference. Extensive experiments show Crab$^{+}$ covers broader tasks than existing unified models while outperforming specialized models on various benchmarks. We successfully reverse the negative transfer trend, achieving positive transfer where multi-task learning surpasses single-task baselines in nearly 88% of tasks. These results hold across diverse AV-LLM paradigms and are validated through in-depth visualization, positioning Crab$^{+}$ as a robust step towards holistic audio-visual scene understanding.

音视频理解多任务学习大模型正迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。