用音频语言模型修复语音分离中的错误,提升真实场景下的分离效果。
SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation
- 先分离后在文本域用音频语言模型纠错并重合成语音
- 在复杂声学环境下显著降低分离误差,提升语音清晰度
- 适合需要高鲁棒性语音处理的工业应用
当前语音分离技术虽能处理长段混合音频,但在噪声和混响等真实环境下游明显受限,常导致分离语音出现失真或伪影。为此,我们提出SepALM,一种创新方法:利用音频语言模型(ALM)在初步分离后,于文本域对语音进行纠错与重合成。SepALM包含分离器、校正器、合成器与对齐器四部分。通过端到端的ALM误差修正机制,有效避免错误累积,并规避传统ASR+LLM融合方法中的优化难题。我们还引入链式思维(CoT)提示与知识蒸馏技术,促进ALM推理与训练。实验表明,SepALM不仅提升分离精度,更显著增强在新声学环境下的适应能力。
原文摘要 · Abstract (English)
While contemporary speech separation technologies adeptly process lengthy mixed audio waveforms, they are frequently challenged by the intricacies of real-world environments, including noisy and reverberant settings, which can result in artifacts or distortions in the separated speech. To overcome these limitations, we introduce SepALM, a pioneering approach that employs audio language models (ALMs) to rectify and re-synthesize speech within the text domain following preliminary separation. SepALM comprises four core components: a separator, a corrector, a synthesizer, and an aligner. By integrating an ALM-based end-to-end error correction mechanism, we mitigate the risk of error accumulation and circumvent the optimization hurdles typically encountered in conventional methods that amalgamate automatic speech recognition (ASR) with large language models (LLMs). Additionally, we have developed Chain-of-Thought (CoT) prompting and knowledge distillation techniques to facilitate the reasoning and training processes of the ALM. Our experiments substantiate that SepALM not only elevates the precision of speech separation but also markedly bolsters adaptability in novel acoustic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。