arXiv:2501.00039eess.AScs.CL2025-01中稿 · ICASSP 2025被引 5

用强化学习让大模型更好听懂有障碍的语音

Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning

  • 把语音转成模型可处理的音频令牌,再用带语义奖励的强化学习微调
  • 在不同场景下适应障碍语音时,强化学习比传统微调效果好得多
  • 适合需要适配特殊语音的场景,如残障人士或口音复杂用户

我们提出一种能处理语音输入的大语言模型(LLM),并通过基于人类偏好的强化学习(RLHF)进一步优化,使其在识别障碍语音方面优于传统微调方法。该方法将LLM词汇表中的低频文本标记替换为音频标记,并通过带文本转录的语音数据进行微调,使模型具备语音识别能力。随后,采用基于句法和语义准确率的强化学习策略,提升模型对障碍语音的泛化能力。尽管该模型在标准语音识别任务上未超越现有系统,但在不同环境下的障碍语音适应中,使用自定义奖励的强化学习微调显著优于监督微调,展现出一种有前景的大模型语音识别优化路径。

原文摘要 · Abstract (English)

We introduce a large language model (LLM) capable of processing speech inputs and show that tuning it further with reinforcement learning on human preference (RLHF) enables it to adapt better to disordered speech than traditional fine-tuning. Our method replaces low-frequency text tokens in an LLM's vocabulary with audio tokens and enables the model to recognize speech by fine-tuning it on speech with transcripts. We then use RL with rewards based on syntactic and semantic accuracy measures generalizing the LLM further to recognize disordered speech. While the resulting LLM does not outperform existing systems for speech recognition, we find that tuning with reinforcement learning using custom rewards leads to substantially better performance than supervised fine-tuning of the language model, specifically when adapting to speech in a different setting. This presents a compelling alternative tuning strategy for speech recognition using large language models.

语音识别大模型强化学习障碍语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。