arXiv:2509.18377cs.CL2025-09被引 2

让会议中用户一键修正说话人识别错误,提升语音转录准确率。

Interactive In-Meeting Speaker Correction with Human Feedback

  • 用大模型生成摘要帮助用户快速发现说话人错误
  • 用户反馈后实时更新转录结果并新增说话人特征
  • 在AMI数据集上将错误率降低32%,说话人混淆减少53%

大多数自动语音处理系统以开环模式运行,缺乏用户对谁说了什么的反馈。我们提出一种基于大模型的会议内说话人纠错系统,允许用户通过简短纠正反馈修复说话人识别错误。系统先进行流式语音识别与说话人分离,再由大模型生成简洁摘要,辅助用户识别关键说话人错误,并通过更新带说话人标注的转录文本和在线注册说话人特征来融入用户反馈。为应对语音处理、大模型分析和用户反馈中的误差,我们设计了多种机制以更精确识别用户意图。此外,构建了大模型驱动的用户反馈模拟器,实现工作流可复现性与大规模评估。在AMI头戴麦克风测试集上,相比流式基线(Google ASR + ECAPA),本系统将总错误率降低31.99%,说话人误标错误减少52.68%。

原文摘要 · Abstract (English)

Most automatic speech processing systems operate in ``open loop'' mode without user feedback about who said what, yet human-in-the-loop workflows can potentially enable higher accuracy. We propose an LLM-assisted in-meeting speaker correction system that lets users fix speaker attribution errors through brief corrective feedback. After performing streaming ASR and diarization, the system presents concise LLM-generated summaries to help users identify important speaker errors, and it incorporates user feedback by updating the speaker-attributed transcript and adding online speaker enrollments. To make this workflow effective despite errors in speech processing, LLM analysis, and user feedback, we developed several mechanisms to identify the intended correction more precisely. Further, we built an LLM-driven user feedback simulation to evaluate the workflow reprodubilty and at scale. Applied to the AMI headset test set, our system substantially reduces the DER from a streaming baseline (Google ASR + ECAPA) by 31.99% and speaker substitution error by 52.68%.

说话人分离大模型应用会议转录人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。