测试音频大模型是否真听懂声音,而非只靠文字猜答案。
DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models
- 设计2700+冲突样本,测试情感语调、背景音、说话人身份三维度
- 通过逐步增强文字干扰,发现模型仍以文本为主导
- 新指标量化模型对声音与文字的依赖程度,适合评测音频模型
近期音频多模态大语言模型在语音基准上表现优异,但其是否真正理解声学信号仍不明确。为此,我们提出DEAF(诊断性评估声学忠实度)基准,包含超过2700个跨三个声学维度(情感语调、背景音、说话人身份)的冲突刺激样本。我们设计了一个控制性的多层级评估框架,逐步增加文本影响,从内容语义冲突到误导性提示及其组合,以分离内容驱动偏差与提示诱导迎合。我们进一步引入诊断指标,量化模型对文本线索相对于声学信号的依赖程度。对七种音频多模态大模型的评估显示,存在一致的文本主导现象:模型虽能感知声学变化,但预测主要由文本输入驱动,揭示了标准语音基准高分与真实声学理解之间的差距。
原文摘要 · Abstract (English)
Recent Audio Multimodal Large Language Models (Audio MLLMs) demonstrate impressive performance on speech benchmarks, yet it remains unclear whether these models genuinely process acoustic signals or rely on text-based semantic inference. To systematically study this question, we introduce DEAF (Diagnostic Evaluation of Acoustic Faithfulness), a benchmark of over 2,700 conflict stimuli spanning three acoustic dimensions: emotional prosody, background sounds, and speaker identity. Then, we design a controlled multi-level evaluation framework that progressively increases textual influence, ranging from semantic conflicts in the content to misleading prompts and their combination, allowing us to disentangle content-driven bias from prompt-induced sycophancy. We further introduce diagnostic metrics to quantify model reliance on textual cues over acoustic signals. Our evaluation of seven Audio MLLMs reveals a consistent pattern of text dominance: models are sensitive to acoustic variations, yet predictions are predominantly driven by textual inputs, revealing a gap between high performance on standard speech benchmarks and genuine acoustic understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。