用语音信息增强大模型,精准修正语音识别中的说话人错误
SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models
- 用音频特征条件化大语言模型,提升纠错上下文精度
- 在多个数据集上降低24%-43%的说话人错误率
- 无需复杂后处理,简单解码策略减少幻觉
说话人分离(SD)是现代端到端自动语音识别系统的关键组件。传统基于音频的独立SD系统在说话人转换和重叠语音场景下常引入说话人错误。近期研究表明,通过利用转录文本中的词汇上下文,微调的大语言模型可作为第二阶段的说话人错误修正器。本文提出一种新颖的声学条件化方法,将更细粒度的声学分离信息提供给大语言模型。同时,我们证明一种更简单的约束解码策略能有效减少大语言模型的幻觉,且无需复杂后处理。相比第一阶段的声学SD,该方法在Fisher、Callhome和RT03-CTS数据集上显著降低了24%-43%的说话人错误率。
原文摘要 · Abstract (English)
Speaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often introduce speaker errors, particularly during speaker transitions and overlapping speech. Recently, language models including fine-tuned large language models (LLMs) have shown to be effective as a second-pass speaker error corrector by leveraging lexical context in the transcribed output. In this work, we introduce a novel acoustic conditioning approach to provide more fine-grained information from the acoustic diarizer to the LLM. We also show that a simpler constrained decoding strategy reduces LLM hallucinations, while avoiding complicated post-processing. Our approach significantly reduces the speaker error rates by 24-43% across Fisher, Callhome, and RT03-CTS datasets, compared to the first-pass Acoustic SD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。