arXiv:2605.26747cs.AI2026-05

构建了真实场景下的医患语音对话数据集,助力医疗AI对话研究。

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

论文配图:A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks
图 1 · 摘自论文原文
  • 采集机器人与患者、医生与患者的真实对话,覆盖四种疾病。
  • 含111小时语音数据,评测显示Claude Sonnet 4在句选任务中准确率达74.7%。
  • 适合医疗对话、语音理解、大模型评估等方向的研究者使用。

大型语言模型(LLMs)在人工智能领域带来显著提升,但其在文本或语音医疗咨询中的应用仍属开放问题。本文提出MeDial-Speech,一个用于训练和评估医疗AI的新型语音数据集,包含从机器人-患者和医生-患者对话中收集的111+小时真实环境语音数据(未经数据增强),涵盖路易体痴呆、心力衰竭、肩痛和心绞痛四种健康状况。此外,我们设计了一项基于句选的对话评测基准(20个选项),用于评估GPT-5 mini、DeepSeek-V3和Claude Sonnet 4三款先进LLM。实验结果表明,Claude Sonnet 4在句选任务中表现最佳,人工转录下准确率为71.1%,自动转录下达74.7%;且所有模型在预测时均表现出高度自信,无论所选句子是否正确。该数据集可免费用于非商业用途:https://huggingface.co/datasets/hcuayahu/MeDial-Speech。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have brought huge improvements to Artificial Intelligence (AI), which can be applied to general-purpose tasks. However, their application to textual or spoken medical consultations is still an open research problem. This paper proposes MeDial-Speech, a novel speech dataset for training and evaluating Med-AIs that can carry out consultations with patients. It was collected in realistic environments from robot-patient and doctor-patient dialogues, contains 111+ hours of speech data (without data augmentation), and covers four health conditions: Lewy body dementia, heart failure, shoulder pain, and angina. In addition, we propose a dialogue benchmark via sentence selection (with 20 options) to evaluate three state-of-the-art LLMs: GPT-5 mini, DeepSeek-V3, and Claude Sonnet 4. Experimental results reveal that Claude Sonnet 4 is the best in sentence selection, with 71.1% accuracy using manual transcriptions and 74.7% using automatic transcriptions, and that all LLMs are highly overconfident in their probabilistic predictions, regardless of selecting correct or incorrect sentences in medical dialogues. This dataset is free of charge for non-commercial purposes at: https://huggingface.co/datasets/hcuayahu/MeDial-Speech

医疗对话语音数据集大模型评估真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。