测试音频大模型在真实声学环境下的鲁棒性,发现其高阶推理能力易崩溃。
RSA-Bench: Benchmarking Audio Large Models in Real-World Acoustic Scenarios
- 用真实环境音景自然叠加语音,模拟复杂声学场景进行评测
- 低级识别尚可,但高级推理任务在干扰下性能大幅下降
- 去噪处理反而损害模型表现,尤其对语义失真敏感
尽管音频大模型(ALMs)已取得显著进展,但在真实部署中仍显脆弱。现有评估多依赖合成高斯噪声或简单单源干扰,难以捕捉真实物理环境中复杂的多层声学动态——即“声学生态”。为此,我们提出RSA-Bench,一个全面的鲁棒性基准,通过高保真听觉场景模拟来压力测试所有音频大模型。不同于传统方法,我们通过自然叠加多样环境音景(包括牧场、极端天气、教室、户外)到纯净语音信号上,覆盖不同干扰强度。在六项核心任务(从基础感知到复杂推理)上的评估揭示三个宏观洞察:(I) 感知-认知鸿沟:模型在低级识别任务中保持相对鲁棒,但在高压下高阶推理任务出现功能崩溃;(II) 场景敏感性:类人声音干扰(如背景笑声)比机械噪声更具破坏性,挑战模型听觉注意力机制;(III) 去噪悖论:标准语音增强常加剧性能退化,因所有音频大模型对去噪引入的语义失真高度敏感。
原文摘要 · Abstract (English)
While Audio Large Models (ALMs) have achieved remarkable proficiency, their robustness remains brittle in real-world deployment. Existing evaluations largely rely on synthetic Gaussian noise or simplistic single-source interference, failing to capture the intricate, multi-layered acoustic dynamics -- or ``Acoustic Ecology'' -- that characterize authentic physical environments. To bridge this ecological gap, we introduce \textbf{RSA-Bench}, a comprehensive robustness benchmark designed to stress-test ALLMs through high-fidelity auditory scene simulations. Unlike traditional methods, we construct evaluation samples by naturally superimposing diverse environmental soundscapes -- spanning \textit{Pasture}, \textit{Extreme Weather}, \textit{Classroom}, and \textit{Outdoors} -- onto clean speech signals across a spectrum of interference intensities. By evaluating models on six core tasks ranging from fundamental perception to complex reasoning, our study unveils three macro-level insights: \textbf{(I) The Perception-Cognition Gap:} Models maintain relative resilience in low-level recognition but suffer a \textbf{functional collapse} in high-order reasoning tasks under stress; \textbf{(II) Scenario Sensitivity:} ``Vocal-like'' interference (e.g., background laughter) proves significantly more destructive than mechanical noise, challenging the model's auditory attention mechanisms; and \textbf{(III) The Denoising Paradox:} Standard speech enhancement often exacerbates performance degradation, as ALLMs prove highly sensitive to the semantic distortions introduced by denoising artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。