arXiv:2509.15612cs.SDeess.AS2025-09

用思维链和强化学习提升语音场景中目标说话人识别效果

Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition

  • 引入思维链与强化学习训练框架,增强模型推理能力
  • 在真实混叠语音中,目标说话人识别准确率显著提升
  • 适合需要精准区分多人对话的语音分析任务

目标说话人自动语音识别(TS-ASR)旨在从多人混合的鸡尾酒会场景中,识别并转录指定说话人的语音。尽管大型音频语言模型(LALMs)已为该任务带来新进展,但其性能仍有优化空间。由于TS-ASR需深入理解语音信号、区分不同说话人并处理重叠话语,因此特别适合采用推理引导的方法。本文提出一种新框架,将思维链(CoT)与强化学习(RL)训练引入TS-ASR。构建了首个面向TS-ASR的思维链数据集,模型先在常规数据上训练,再在CoT数据上微调,最后使用精选数据进行强化学习训练,以增强泛化推理能力。实验表明,结合CoT与RL训练后,TS-ASR性能显著提升,验证了该方法的有效性。

原文摘要 · Abstract (English)

Target Speaker Automatic Speech Recognition (TS-ASR) aims to transcribe the speech of a specified target speaker from multi-speaker mixtures in cocktail party scenarios. Recent advancement of Large Audio-Language Models (LALMs) has already brought some new insights to TS-ASR. However, significant room for optimization remains for the TS-ASR task within the LALMs architecture. While Chain of Thoughts (CoT) and Reinforcement Learning (RL) have proven effective in certain speech tasks, TS-ASR, which requires the model to deeply comprehend speech signals, differentiate various speakers, and handle overlapping utterances is particularly well-suited to a reasoning-guided approach. Therefore, we propose a novel framework that incorporates CoT and RL training into TS-ASR for performance improvement. A novel CoT dataset of TS-ASR is constructed, and the TS-ASR model is first trained on regular data and then fine-tuned on CoT data. Finally, the model is further trained with RL using selected data to enhance generalized reasoning capabilities. Experiment results show a significant improvement of TS-ASR performance with CoT and RL training, which demonstrates the effectiveness of the proposed CoT and RL training methods adapted for the TS-ASR task.

语音识别思维链强化学习多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。