arXiv:2505.22045cs.MMcs.CV2025-05中稿 · INTERSPEECH 2025被引 2

通过动态抑制误导性视觉信息,提升视频引导音频描述的准确性。

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

  • 用注意力熵分析自动识别并压制错误视觉线索。
  • 在音频-视觉错配场景中,性能显著优于现有方法。
  • 推理速度提升约6倍,适合实时应用。

当前视觉引导的音频描述系统在真实场景(如配音内容或离屏声音)中常忽视音画错位问题。为此,我们提出一种基于熵感知的门控融合框架,通过跨模态不确定性量化动态调节视觉信息流。该方法利用交叉注意力层中的注意力熵分析,自动识别并抑制融合过程中的误导性视觉线索。同时,我们设计了一种批量音画错乱技术,生成合成的错配训练样本,显著增强模型对对齐噪声的鲁棒性。在AudioCaps基准上的评估表明,我们的系统在音画错配场景下性能显著优于现有基线,且推理速度相较基线提升约6倍。

原文摘要 · Abstract (English)

Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge this critical gap, we present an entropy-aware gated fusion framework that dynamically modulates visual information flow through cross-modal uncertainty quantification. Our novel approach employs attention entropy analysis in cross-attention layers to automatically identify and suppress misleading visual cues during modal fusion. Complementing this architecture, we develop a batch-wise audiovisual shuffling technique that generates synthetic mismatched training pairs, greatly enhancing model resilience against alignment noise. Evaluations on the AudioCaps benchmark demonstrate our system's superior performance over existing baselines, especially in mismatched modality scenarios. Furthermore, our solution demonstrates an approximately 6x improvement in inference speed compared to the baseline.

音画对齐多模态生成模型推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。