arXiv:2603.02813eess.AS2026-03被引 5

构建医疗对话理解基准,评估多说话人真实场景下的语音系统性能。

Benchmarking Speech Systems for Frontline Health Conversations: The DISPLACE-M Challenge

  • 针对医护与患者对话设计多说话人语音处理任务
  • 40小时开发数据+15小时盲评数据,使用DER、tcpWER等指标评测
  • 提供语音识别、话题识别等四类任务基线,适合医疗AI研究者参考

DISPLACE-M挑战提出了一项面向真实医疗对话的对话智能基准,聚焦前线医护人员与求诊者之间的目标导向型多说话人互动。这些对话具有自发性、噪声大和语音重叠等特点。挑战方发布了包含40小时开发数据和15小时盲评数据的医学对话数据集。针对四个任务——说话人分离、自动语音识别、话题识别和对话摘要——提供了基线系统,以实现一致评估。系统性能通过说话人分离错误率(DER)、时间约束最小排列词错误率(tcpWER)和ROUGE-L进行衡量。本文描述了第一阶段评估的数据、任务和基线系统,并总结了评估结果。

原文摘要 · Abstract (English)

The DIarization and Speech Processing for LAnguage understanding in Conversational Environments - Medical (DISPLACE-M) challenge introduces a conversational AI benchmark for understanding goal-oriented, real-world medical dialogues. The challenge addresses multi-speaker interactions between frontline health workers and care seekers, characterized by spontaneous, noisy and overlapping speech. As part of the challenge, medical conversational dataset comprising 40 hours of development and 15 hours of blind evaluation recordings was released. We provided baseline systems across 4 tasks - speaker diarization, automatic speech recognition, topic identification and dialogue summarization - to enable consistent benchmarking. System performance is evaluated using diarization error rate (DER), time-constrained minimum-permutation word error rate (tcpWER) and ROUGE-L. This paper describes the Phase-I evaluation - data, tasks and baseline systems - along with the summary of the evaluation results.

医疗AI语音识别对话系统多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。