发现音频模型常被文本误导,提出修复方案提升判断准确性。
Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models

- 通过保留音频、移除冲突文本,检测模型真实偏好
- 64.1%样本中音频应答优先于文本,说明音频信息被覆盖
- 无需训练的解码规则可提升精度,且跨模态适用
音频-语言模型(ALMs)常因文本冲突而忽略清晰的音频证据。我们通过固定音频、仅移除冲突文本的反事实实验,考察模型是否具备音频支持答案的能力。在五个ALMs和四个冲突任务中,64.1%的样本出现偏好反转:相同音频分支更倾向音频支持的答案,而联合分支则偏向文本支持的答案。这表明音频信息已被编码,但在仲裁中被覆盖。激活修补进一步定位到答案位置计算阶段,修补效果与输出候选得分差高度相关(斯皮尔曼秩相关=0.93)。基于此诊断,我们提出无训练解码规则GACL,通过插值联合与同音频得分实现修复。在严格5个百分点忠实度下降预算下,GACL相比最优对比基线提升nAUC 17.8点,并在视觉-文本仲裁中无需微调即实现最高+40.5点性能增益。
原文摘要 · Abstract (English)
Audio-language models (ALMs) often follow text that conflicts with audio, even when the audio evidence is clear. This raises a basic question: is the audio-supported answer unavailable, or is it represented but overridden by the conflicting text? We examine this question using a same-audio counterfactual that keeps the audio fixed, removes only the conflicting text, and measures the resulting shift in model preference. Across five ALMs and four conflict tasks, 64.1% of conflict samples show a sign flip: the same-audio branch prefers the audio-supported answer, whereas the joint branch prefers the text-supported answer. This pattern suggests that the relevant audio evidence is encoded but loses in arbitration. Activation patching further localizes the reversal to answer-position computation, and patching effects closely track output candidate-score differences (Spearman rho=0.93). Using this diagnostic, we propose Gated Audio Counterfactual Logit Correction (GACL), a training-free decoding rule that interpolates between joint and same-audio scores. Under a strict 5 pp faithfulness-drop budget, GACL improves nAUC by 17.8 points over the best contrastive baseline and transfers without retuning to vision-text arbitration (up to +40.5 pp).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。