无需微调,插件式降噪提升语音模型在嘈杂环境下的表现
Focus Then Listen: An Empirical Study of Plug-and-Play Audio Enhancer for Noise-Robust Large Audio Language Models
- 先分离语音与非语音信号,再根据指令选择目标模态
- 在多种噪声条件下提升多个语音模型的识别准确率
- 适合需要快速部署且不依赖特定噪声数据的场景
大型音频语言模型(LALMs)是用于音频理解的基础模型。现有LALMs在真实嘈杂环境中性能显著下降,因语音与非语音声音干扰。虽然噪声感知微调可增强鲁棒性,但需任务特定噪声数据并进行昂贵重训练,限制可扩展性。为此,我们提出即插即用的降噪模块Focus-Then-Listen(FTL),通过将输入波形分离为语音与非语音成分,利用模态路由根据用户指令预测目标模态(如语音),再通过模态感知融合块生成任务自适应增强信号,从而提升下游感知与推理能力。多模型、多任务实验表明,FTL在无需对LALMs进行微调的情况下,有效提升了不同噪声水平下的性能。
原文摘要 · Abstract (English)
Large audio language models (LALMs) are a class of foundation models for audio understanding. Existing LALMs tend to degrade significantly in real-world noisy acoustic conditions where speech and non-speech sounds interfere. While noise-aware fine-tuning can improve robustness, it requires task-specific noisy data and expensive retraining, limiting scalability. To address this issue, we propose Focus-Then-Listen (FTL), a plug-and-play audio enhancer that improves LALMs' noise robustness. Specifically, FTL first separates the input waveform into speech and non-speech, and a modality router is applied to predict the target audio modality (e.g., speech) based on the user's instruction. Finally, a modality-aware fusion block generates a task-adaptive enhanced signal for improved downstream perception and reasoning. Experiments across multiple LALMs and tasks show that FTL improves performance across different noise levels without fine-tuning on LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。