用多轮LLM迭代提升法语临床对话转写与说话人分离准确率
Iterative LLM-based improvement for French Clinical Interview Transcription and Speaker Diarization
- 通过交替进行说话人识别和词语识别的多轮处理流程优化转写
- 自杀预防通话中词错误率显著降低,清醒神经外科术前咨询稳定有效
- 适合需要高精度语音转写的医疗场景,可离线部署
法语临床对话的自动语音识别仍具挑战性,自然口语中的词错误率常超过30%。本文提出一种多轮LLM后处理架构,交替执行说话人识别与词语识别步骤,以提升转写准确性和说话人归属。在两个法语临床数据集(自杀预防电话咨询和清醒神经外科术前会诊)上开展消融实验,考察模型选择、提示策略、轮次顺序和迭代深度四项设计。使用Qwen3-Next-80B模型,在自杀预防对话中经威尔科克森符号秩检验验证词错误率显著下降(p<0.05, n=18),在清醒神经外科会诊中保持稳定(n=10),无输出失败,计算成本可控(RTF 0.32),表明该方法具备离线临床部署可行性,有待更大规模语料验证。
原文摘要 · Abstract (English)
Automatic speech recognition for French medical conversations remains challenging, with word error rates often exceeding 30% in spontaneous clinical speech. This study proposes a multi-pass LLM post-processing architecture alternating between Speaker Recognition and Word Recognition passes to improve transcription accuracy and speaker attribution. Ablation studies on two French clinical datasets (suicide prevention telephone counseling and preoperative awake neurosurgery consultations) investigate four design choices: model selection, prompting strategy, pass ordering, and iteration depth. Using Qwen3-Next-80B, Wilcoxon signed-rank tests confirm significant WDER reductions on suicide prevention conversations (p<0.05, n=18), while maintaining stability on awake neurosurgery consultations (n=10), with zero output failures and acceptable computational cost (RTF 0.32), suggesting feasibility for offline clinical deployment, pending validation on larger corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。