arXiv:2509.15082eess.AS2025-09被引 2

用大模型让语音识别自动知道谁说了什么、是谁,无需训练就提升准确率。

From Who Said What to Who They Are: Modular Training-free Identity-Aware LLM Refinement of Speaker Diarization

  • 用大模型对语音分段和识别结果做语义关联,自动修正说话人标签
  • 在真实医患对话数据上,错误率比基线降低29.7%
  • 无需训练即可部署,适合医疗、访谈等需身份识别的场景

说话人分离(SD)在真实动态环境中仍具挑战性,且常与自动语音识别(ASR)联合使用。现有非模块化框架缺乏灵活性,无法提供真实说话人身份。本文提出一种无需训练的模块化流程,结合现成的SD、ASR与大语言模型(LLM),通过结构化提示对齐输出结果,利用对话上下文语义连续性来修正低置信度说话人标签,并赋予角色身份,同时合并被拆分的同一说话人。在真实患者-医生对话数据集上,该方法相比基线重构后的SD+ASR系统,相对错误率降低29.7%。该方案无需额外训练即可显著提升分段与身份识别性能,实现完整的端到端语音分析流水线。

原文摘要 · Abstract (English)

Speaker diarization (SD) remains challenging in real-world scenarios due to dynamic environments and unknown speaker numbers. SD is rarely used alone and is typically paired with automatic speech recognition (ASR). However, existing non-modular SD+ASR frameworks lack flexibility and do not provide true speaker identities. We propose a training-free modular pipeline combining off-the-shelf SD, ASR, and a large language model (LLM) to determine who spoke, what was said, and who they are. Using structured LLM prompting on reconciled SD and ASR outputs, our method leverages semantic continuity in conversational context to refine low-confidence speaker labels and assigns role identities while correcting split speakers. On a real-world patient-clinician dataset, our approach achieves a 29.7% relative error reduction over baseline reconciled SD and ASR. It enhances diarization performance without additional training and delivers a complete pipeline for SD, ASR, and speaker identity detection in practical applications.

说话人分离大模型应用身份识别医疗语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。