arXiv:2601.04658cs.SDcs.AI2026-01中稿 · ICASSP 2026被引 2

用柯西-施瓦茨散度对齐音视频特征,提升大模型生成音频描述质量

LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence

  • 通过最小化柯西-施瓦茨散度实现音视频嵌入的跨模态对齐
  • 在AudioCaps上达到最新最优性能,显著提升生成准确性
  • 适合做多模态理解与大模型应用的研究者参考

自动音频描述旨在解析输入音频的语义内容。近期工作采用大语言模型(LLM)作为文本解码器,以利用其推理能力。然而,先前方法在将音频特征投影到LLM嵌入空间时未考虑跨模态对齐,未能充分释放其潜力。为此,我们提出LAMB框架,通过柯西-施瓦茨散度桥接音频嵌入与LLM文本嵌入空间之间的模态鸿沟。LAMB引入跨模态对齐器,在最大化互信息的同时最小化柯西-施瓦茨散度,实现全局与词级别音视频嵌入的紧密对齐。进一步设计双流适配器提取富含语义的音频嵌入,为对齐器提供更丰富信息。最终,基于对齐后的音频嵌入,提出的词级引导模块直接在LLM文本嵌入空间计算得分,引导生成字幕的输出逻辑。实验表明,该框架增强了LLM解码器的推理能力,在AudioCaps数据集上取得当前最优表现。

原文摘要 · Abstract (English)

Automated Audio Captioning aims to describe the semantic content of input audio. Recent works have employed large language models (LLMs) as a text decoder to leverage their reasoning capabilities. However, prior approaches that project audio features into the LLM embedding space without considering cross-modal alignment fail to fully utilize these capabilities. To address this, we propose LAMB, an LLM-based audio captioning framework that bridges the modality gap between audio embeddings and the LLM text embedding space. LAMB incorporates a Cross-Modal Aligner that minimizes Cauchy-Schwarz divergence while maximizing mutual information, yielding tighter alignment between audio and text at both global and token levels. We further design a Two-Stream Adapter that extracts semantically enriched audio embeddings, thereby delivering richer information to the Cross-Modal Aligner. Finally, leveraging the aligned audio embeddings, a proposed Token Guide directly computes scores within the LLM text embedding space to steer the output logits of generated captions. Experimental results confirm that our framework strengthens the reasoning capabilities of the LLM decoder, achieving state-of-the-art performance on AudioCaps.

音频描述大模型跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。