用测试时计算提升语音大模型在嘈杂环境下的听觉认知能力
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
- 引入测试时计算方法,增强语音大模型推理时的认知能力
- 在复杂听觉任务中,模型性能显著提升,噪声下表现更稳定
- 适合开发助听设备、语音助手等实际应用的智能系统
大型语言模型(LLMs)在自然语言处理中表现出色,推动了音频大语言模型(Audio LLMs)在语音处理中的多模态拓展。尽管在语音识别与合成方面表现优异,但其在真实环境中面临的听觉认知挑战——如背景噪声或重叠语音下的音频理解与记忆——仍不明确。由于缺乏能模拟真实听觉场景的多样化数据集及标注困难,音频大模型的再训练受限。而测试时计算(TTC)方法虽可提升文本类大模型的推理能力,但如何有效应用于音频大模型尚不清晰。本研究通过自建数据库,评估五种音频大模型的听觉认知能力,并提出五种TTC方法以增强其推理阶段表现。结果表明,模型在高难度听觉任务中性能下降,而所提方法显著提升了其听觉认知能力,为助听设备、语音助手等实际应用中的自适应、鲁棒音频大模型发展提供支持。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown exceptional versatility in natural language processing, prompting recent efforts to extend their multimodal capabilities to speech processing through the development of audio large language models (Audio LLMs). While Audio LLMs excel in tasks such as speech recognition and synthesis, it remains unclear how they perform when faced with the auditory cognitive challenges posed by real-world environments, such as audio comprehension and listening recall, particularly in the presence of background noise or overlapping speech. Unlike text-based LLMs, which have access to vast amounts of text data for pre-training, retraining Audio LLMs with diverse auditory cognitive scenes is difficult due to the limited datasets that simulate real-world auditory cognitive scenarios and the challenge of acquiring auditory cognitive labels for training. While test-time compute (TTC) methods have been shown to enhance the capabilities of text-based LLMs during inference, a key challenge lies in designing these TTC methods to improve the auditory capabilities of Audio LLMs. This study aims to address these two research gaps by: i) exploring the auditory cognitive capabilities of Audio LLMs, and ii) enhancing their capabilities using TTC approaches. We have investigated five different Audio LLMs for auditory cognition using a \textit{self-collected} database and have proposed five TTC approaches to enhance auditory cognitive capabilities during inference. Our findings reveal that Audio LLMs performance decreases in more challenging auditory cognitive tasks. The proposed TTC approaches significantly enhance cognitive auditory capabilities, advancing the development of more adaptable and resilient Audio LLMs for practical applications such as assistive listening devices, voice-based AI assistants, and communication technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。