首个面向中文医疗对话的实时双工语音识别基准,支持多轮交互与重叠说话人
MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
- 构建真实医疗场景下的双工语音数据集,含5805个标注会话
- 提出流式分割与角色追踪方法,实现长上下文语音识别与对话记忆
- 适用于医疗AI助手研发者与语音识别研究者评估双工系统性能
临床对话中的自动语音识别(ASR)需要应对全双工交互、说话人重叠和低延迟约束,但现有公开基准仍匮乏。本文提出MMedFD,首个面向真实中文医疗场景的多轮全双工语音识别语料库。该数据来自已部署的AI助手机器人,包含5,805个经标注的会话,同步提供用户与混合通道视角、RTTM/CTM时间标记及角色标签。我们设计了一种模型无关的流水线,用于流式分割、说话人归属与对话记忆建模,并对Whisper-small在角色拼接音频上进行微调以实现长上下文识别。评估包含WER、CER与HC-WER(衡量医疗概念级准确率)。大语言模型生成回复通过评分制与成对比较协议评估。MMedFD建立可复现的框架,用于医疗部署中流式ASR与端到端双工代理的基准测试。数据集及相关资源已公开:https://github.com/Kinetics-JOJO/MMedFD
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) in clinical dialogue demands robustness to full-duplex interaction, speaker overlap, and low-latency constraints, yet open benchmarks remain scarce. We present MMedFD, the first real-world Chinese healthcare ASR corpus designed for multi-turn, full-duplex settings. Captured from a deployed AI assistant, the dataset comprises 5,805 annotated sessions with synchronized user and mixed-channel views, RTTM/CTM timing, and role labels. We introduce a model-agnostic pipeline for streaming segmentation, speaker attribution, and dialogue memory, and fine-tune Whisper-small on role-concatenated audio for long-context recognition. ASR evaluation includes WER, CER, and HC-WER, which measures concept-level accuracy across healthcare settings. LLM-generated responses are assessed using rubric-based and pairwise protocols. MMedFD establishes a reproducible framework for benchmarking streaming ASR and end-to-end duplex agents in healthcare deployment. The dataset and related resources are publicly available at https://github.com/Kinetics-JOJO/MMedFD
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。