评测大模型对音频降质的感知能力,发现现有模型表现不佳。
MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

- 设计多轮多音频对话形式,考察模型对降质类型、程度的识别与推理。
- 18个主流模型在降质诊断和比较上普遍表现不佳,准确率偏低。
- 适合研究音频-语言模型鲁棒性或听觉认知的学者使用。
大音频-语言模型(LALMs)在语音、音乐和一般声音事件理解方面取得显著进展,但其对音频信号退化机制的理解仍不充分。现有基准主要评估语义理解、事件识别或高层音频推理,却未回答一个基本问题:LALMs是否能感知音频质量差异?我们提出MRMAD——一个多轮多音频退化评估基准,用于评测LALMs在语音、音乐和声音场景中对音频退化的感知与理解能力。该基准以多轮对话形式呈现多个音频输入,要求模型识别退化类型、比较严重程度,并在多轮交互中持续追踪退化变化。与现有单轮基准不同,MRMAD考察模型能否在新证据下维持一致的退化假设,并理解低层声学现象。通过对18个代表性LALMs(从非思考到推理及全模态模型)的系统评估,我们发现当前模型虽能识别粗粒度内容,但无法可靠地诊断、比较或推理退化问题。人类评估进一步揭示了模型与人类听者之间的显著感知差距。MRMAD暴露了音频-语言理解中一个被忽视的关键维度,并为构建更适应真实声学环境的未来LALMs提供了诊断基础。
原文摘要 · Abstract (English)
Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event recognition, or high-level audio reasoning, leaving a basic question unanswered: Do LALMs understand the differences in audio quality? We introduce MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating audio degradation perception and understanding in LALMs. MRMAD spans speech, music, and sound, and frames evaluation as multi-turn dialogues across multiple audio inputs, requiring models to identify types of degradation, compare severity, and perceive corruption changes across turns. Unlike current single-turn audio-language benchmarks, MRMAD evaluates whether LALMs can maintain consistent degradation hypotheses with new evidence and comprehend low-level acoustic phenomena over multi-turn dialogues. Through a systematic evaluation of 18 representative LALMs from non-thinking to reasoning and Omni models, we find that current models often recognize coarse content while failing to diagnose, compare, or reason about degradations reliably. Human evaluations further reveal a significant perception gap between LALMs and human listeners. MRMAD thus exposes a critical yet overlooked aspect of audio-language understanding and provides a diagnostic foundation for building future LALMs that are robust to real-world acoustic conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。