解决医生患者对话中混用印地语和英语的医疗信息提取难题
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
- 采用基于向量聚类的端到端说话人分离技术,精准处理密集重叠语音
- 在真实医疗对话数据集上实现18.59%的词错误率,优于多数基线方法
- 开源系统在25个参赛者中排名第一,适合医疗语音处理研究者参考
从混用印地语与英语的临床口语对话中提取患者医疗状况极具挑战,因存在快速换人和高度重叠语音。我们在真实世界Hinglish医疗对话数据集DISPLACE-M上评估了一个鲁棒系统。提出端到端神经说话人分离结合向量聚类方法(EEND-VC),有效解决医生-患者对话中的密集重叠问题。在语音识别方面,通过领域特定微调、天城文脚本规范化及对话级大模型纠错,将Qwen3 ASR模型应用于该场景,实现18.59%的词错误率(tcpWER)。对比开放与专有大模型在医疗状况提取上的表现,我们构建的文本级级联系统与多模态端到端音频框架进行基准测试。尽管专有端到端模型性能达到上限,但我们的开源级联架构仍具竞争力,在DISPLACE-M挑战赛中位列25名参赛者首位。所有代码与实现均已公开。
原文摘要 · Abstract (English)
Extracting patient medical conditions from code-switched clinical spoken dialogues is challenging due to rapid turn-taking and highly overlapped speech. We present a robust system evaluated on the DISPLACE-M dataset of real-world Hinglish medical conversations. We propose an End-to-End Neural Diarization with Vector Clustering approach (EEND-VC) to accurately resolve dense and speaker overlaps in Doctor-Patient Conversations (DoPaCo). For transcription, we adapt a Qwen3 ASR model via domain-specific fine-tuning, Devanagari script normalization, and dialogue-level LLM error correction, achieving an 18.59% tcpWER. We benchmark open and proprietary LLMs on medical condition extraction, comparing our text-based cascade system against a multimodal End-to-End (E2E) audio framework. While proprietary E2E models set the performance ceiling, our open cascaded architecture is highly competitive, as it achieved first place out of 25 participants in the DISPLACE-M challenge. All implementations are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。