arXiv:2509.22363cs.LGeess.AS2025-09中稿 · Interspeech 2026被引 3

提出评估音频语言模型推理忠实性的系统框架,揭示其与音频输入的脱节问题。

Investigating Faithfulness in Large Audio Language Models

  • 构建基于音频与思维链干预的评测基准,量化推理与音频的关联性。
  • 发现模型推理虽匹配最终答案,但常脱离音频实际内容,易受干扰误导。
  • 定义无幻觉、整体感知、专注聆听三标准,为可信多模态推理提供依据。

大型音频语言模型(LALMs)通过整合音频编码器与预训练大语言模型,实现复杂的多模态推理任务。尽管这些模型能生成思维链(CoT)解释,但其推理过程的忠实性仍不明确。本文提出一种系统性框架,评估LALMs在输入音频和最终预测之间的推理忠实性。定义了三个音频忠实性标准:无幻觉、整体感知和专注聆听。同时引入一个基于音频和CoT干预的基准测试,用于评估模型表现。实验在Audio Flamingo 3和Qwen2.5-Omni上进行,发现潜在的多模态脱节现象:推理结果虽与最终预测一致,但并不总是强依赖于音频输入,且容易受到幻觉或对抗性扰动的影响。评测接口与结果详见https://poonehmousavi.github.io/faithfulness/。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) integrate audio encoders with pretrained Large Language Models to perform complex multimodal reasoning tasks. While these models can generate Chain-of-Thought (CoT) explanations, the faithfulness of these reasoning chains remains unclear. In this work, we propose a systematic framework to evaluate CoT faithfulness in LALMs with respect to both the input audio and the final model prediction. We define three criteria for audio faithfulness: hallucination-free, holistic, and attentive listening. We also introduce a benchmark based on both audio and CoT interventions to assess faithfulness\footnote{The benchmarking interface and evaluation results are available at https://poonehmousavi.github.io/faithfulness/. Experiments on Audio Flamingo 3 and Qwen2.5-Omni suggest a potential multimodal disconnect: reasoning often aligns with the final prediction but is not always strongly grounded in the audio and can be vulnerable to hallucinations or adversarial perturbations.

音频模型多模态推理可信性思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。