无关音频会干扰大模型文本推理,连沉默都可能让输出变乱。
When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models
- 测试不同无关音频对文本推理的影响,发现干扰普遍存在。
- 沉默和合成噪声均导致准确率下降,持续时间越长影响越严重。
- 自一致性方法可提升稳定性,但需更多计算资源。
大型音频-语言模型(LALMs)融合语音与文本处理,但在嘈杂现实场景中的鲁棒性仍待研究。本文考察了无关音频(如沉默、合成噪声、环境声)对无需音频的文本推理任务的影响。在三个文本基准上,即使非信息性音频也会降低准确率并增加预测波动;干扰程度随音频时长增加、音量升高及解码温度提高而加剧。令人意外的是,沉默的干扰力与合成噪声相当。尽管更大模型更具韧性,但所有系统仍存在漏洞。我们测试了缓解策略,发现提示工程效果有限,而自一致性方法虽能提升稳定性,但计算开销上升。结果揭示跨模态干扰是关键鲁棒性挑战,亟需高效融合机制以应对无关输入。
原文摘要 · Abstract (English)
Large audio-language models (LALMs) unify speech and text processing, but their robustness in noisy real-world settings remains underexplored. We investigate how irrelevant audio, such as silence, synthetic noise, and environmental sounds, affects text reasoning tasks where audio is unnecessary. Across three text-based benchmarks, we find that even non-informative audio reduces accuracy and increases prediction volatility; the severity of interference scales with longer durations, higher amplitudes, and elevated decoding temperatures. Silence, often assumed neutral, destabilizes outputs as strongly as synthetic noise. While larger models show greater resilience, vulnerabilities persist across all evaluated systems. We further test mitigation strategies and find that prompting shows limited effectiveness, whereas self-consistency improves stability at the cost of increased computation. Our results reveal cross-modal interference as a key robustness challenge and highlight the need for efficient fusion strategies that preserve reasoning performance in the presence of irrelevant inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。