构建首个英文医学问诊对话数据集,支持精准病史采集。
MediTOD: An English Dialogue Dataset for Medical History Taking with Comprehensive Annotations
- 基于问卷设计标注体系,覆盖症状、发病、进展等属性
- 包含高精度医学槽位标注,支持诊断级语义理解
- 适合医疗对话、生物医学语言模型研究者使用
医疗任务导向对话系统可辅助医生采集患者病史,助力诊断与治疗决策,缓解医生工作负担并扩大医疗服务可及性。然而,医患对话数据集因隐私限制难以获取,现有数据缺乏涵盖症状及其发生、进展、严重程度等属性的全面标注,且多数非英文,限制了研究社区的使用。为此,我们提出MediTOD,一个面向病史采集任务的英文医患对话数据集。通过与医生合作,设计适配医学领域的问卷式标注方案,由专业医疗人员进行高质量标注,完整记录医学槽位及其属性。我们在监督学习和少样本场景下为自然语言理解、策略学习和自然语言生成任务建立基准,评估来自对话系统与生物医学领域的模型性能。MediTOD已公开,供后续研究使用。
原文摘要 · Abstract (English)
Medical task-oriented dialogue systems can assist doctors by collecting patient medical history, aiding in diagnosis, or guiding treatment selection, thereby reducing doctor burnout and expanding access to medical services. However, doctor-patient dialogue datasets are not readily available, primarily due to privacy regulations. Moreover, existing datasets lack comprehensive annotations involving medical slots and their different attributes, such as symptoms and their onset, progression, and severity. These comprehensive annotations are crucial for accurate diagnosis. Finally, most existing datasets are non-English, limiting their utility for the larger research community. In response, we introduce MediTOD, a new dataset of doctor-patient dialogues in English for the medical history-taking task. Collaborating with doctors, we devise a questionnaire-based labeling scheme tailored to the medical domain. Then, medical professionals create the dataset with high-quality comprehensive annotations, capturing medical slots and their attributes. We establish benchmarks in supervised and few-shot settings on MediTOD for natural language understanding, policy learning, and natural language generation subtasks, evaluating models from both TOD and biomedical domains. We make MediTOD publicly available for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。