提出新方法量化噪声对大模型语音输入的影响,提升实际部署鲁棒性。
SEE: Signal Embedding Energy for Quantifying Noise Interference in Large Audio Language Models
- 基于模型内部表征的结构化激活子空间,评估噪声干扰强度
- 该指标与模型性能相关性高达0.98,可精准反映噪声影响
- 发现传统降噪方法可能适得其反,适合优化语音大模型鲁棒性研究者
大型语音语言模型(LALMs)广泛应用于车载助手、在线会议理解等实时场景。然而,音频输入常受设备与环境噪声污染,导致性能下降。现有研究缺乏对噪声影响的定量分析,主要依赖直觉与经验观察,难以评估实际鲁棒性。为此,本文提出信号嵌入能量(SEE),一种量化噪声强度对LALM输入影响的方法,可区分模型在真实部署中的鲁棒性差异。SEE基于模型内部表示的结构化激活子空间,比原始音频特征更准确捕捉模型对噪声的感知。实验表明,SEE与模型性能高度相关,相关系数达0.98。令人意外的是,传统语音降噪方法对LALMs仅略有改善,甚至在某些情况下增加SEE值并降低性能,表明语音主导的降噪目标与现代LALMs的噪声敏感性存在不匹配。因此,本文基于SEE提出新型输入降噪策略,优于现有方法。本工作为提升LALMs在真实场景下的鲁棒性提供了新度量与优化方向。
原文摘要 · Abstract (English)
Large Audio Language Models (LALMs) have been widely applied in real-time scenarios, such as in-car assistants and online meeting comprehension. In practice, audio inputs are often corrupted by device and environmental noise, leading to performance degradation. However, existing LALM studies on noise lack quantitative analysis and rely mainly on intuition and empirical observation, thus failing to understand practical robustness. To address this issue, we introduce Signal Embedding Energy (SEE), a method for quantifying the impact of noise intensity on LALM inputs, enabling the differentiation of LALM robustness in real-world deployments. SEE introduces a perspective based on structured activation subspaces derived from the model's internal representations, which more accurately captures its perception of noise than raw audio features. Across experiments, SEE exhibits a strong correlation with LALM performance, achieving a correlation of 0.98. Surprisingly, traditional audio denoising methods are only marginally effective for LALMs, and, in some cases, even increase SEE and impair performance. This suggests a mismatch between speech-centric denoising objectives and the noise sensitivity of modern LALMs. Therefore, we propose a mitigation strategy derived from SEE to denoise LALM inputs, outperforming existing denoising methods. This paper introduces a novel metric for noise quantification in LALMs, providing guidance for robustness improvements in real-world deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。