arXiv:2508.01659cs.SDeess.AS2025-08

提出音频共性描述任务,提升多模态大模型的跨模态理解能力。

From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs

  • 通过捕捉多段音频的共同语义,替代差异描述任务。
  • 在多个语音音乐任务上保持更强的通用性能,避免灾难性遗忘。
  • 适合追求跨模态鲁棒性的多模态大模型研究者使用。

音频描述(AC)在多模态大模型(MLLMs)的预训练和微调中对增强音文跨模态理解起关键作用。近期工作提出音频差异描述(ADC),通过多音频输入引导模型描述其差异,以促进细粒度区分。然而,该任务在输入音频事件丰富与输出短句聚焦差异之间产生语义鸿沟,与标准AC任务目标不一致,导致灾难性遗忘。为此,本文提出音频共性描述(ACC),一种挑战相当但更温和的替代方案,引导模型捕捉音频片段间的共享语义而非细节差异。实验表明,ACC不仅提升了在描述基准上的音文理解能力,还在多种语音与音乐任务中更好保留了通用能力,证实其能实现更稳健的跨模态理解,并在泛化性与任务特异性间取得更好平衡。

原文摘要 · Abstract (English)

Audio Captioning (AC) plays a pivotal role in enhancing audio-text cross-modal understanding during the pretraining and finetuning of Multimodal LLMs (MLLMs). To strengthen this alignment, recent works propose Audio Difference Captioning (ADC), which takes multiple audio inputs and encourages the model to describe their differences, thereby promoting fine-grained discrimination. However, despite its effectiveness, ADC introduces a semantic gap between input audios-often rich in diverse events-and the brief, difference-focused short caption. This deviation from AC-style task causes a mismatch with the pretraining objective, leading to catastrophic forgetting. To address this, we propose Audio Commonality Captioning (ACC), a comparably challenging but gentler alternative that guides the model to capture shared semantics across audio clips rather than detailed differences. Experiments show that ACC not only improves audio-text understanding on captioning benchmarks but also better preserves general capabilities across diverse speech and music tasks, confirming its ability to enable more robust cross-modal understanding and achieve a better balance between generalization and task-specific performance in MLLMs.

音频描述跨模态理解多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。