首次研究医患对话中医生意图轨迹,揭示诊疗对话结构规律。
"Where does it hurt?" -- Dataset and Study on Physician Intent Trajectories in Doctor Patient Dialogues
- 基于SOAP框架构建医生意图细粒度分类体系
- 标注超5000轮对话,发现模型难识别症状类别转换
- 成果助力建设辅助诊断系统,适合医疗AI研究者
在医患对话中,医生通过有目的地提问以高效获取信息,实现准确诊断与治疗方案制定。本文首次系统研究医生意图在对话中的演变轨迹。基于‘Ambient Clinical Intelligence Benchmark’(Aci-bench)数据集,我们联合医学专家,依据SOAP框架(主观、客观、评估、计划)构建细粒度意图分类体系,并通过Prolific平台招募大量医学专家,完成超过5000轮对话的标注。该高质量标注数据集为当前医疗意图分类模型提供基准。实验表明,现有模型虽能捕捉对话整体结构,但难以识别SOAP类别间的转换。研究首次揭示常见医患对话结构路径,为设计辅助诊断系统提供关键洞察。此外,我们验证了意图过滤对医疗对话摘要任务的显著提升作用。代码与数据已公开于https://github.com/DATEXIS/medical-intent-classification。
原文摘要 · Abstract (English)
In a doctor-patient dialogue, the primary objective of physicians is to diagnose patients and propose a treatment plan. Medical doctors guide these conversations through targeted questioning to efficiently gather the information required to provide the best possible outcomes for patients. To the best of our knowledge, this is the first work that studies physician intent trajectories in doctor-patient dialogues. We use the `Ambient Clinical Intelligence Benchmark' (Aci-bench) dataset for our study. We collaborate with medical professionals to develop a fine-grained taxonomy of physician intents based on the SOAP framework (Subjective, Objective, Assessment, and Plan). We then conduct a large-scale annotation effort to label over 5000 doctor-patient turns with the help of a large number of medical experts recruited using Prolific, a popular crowd-sourcing platform. This large labeled dataset is an important resource contribution that we use for benchmarking the state-of-the-art generative and encoder models for medical intent classification tasks. Our findings show that our models understand the general structure of medical dialogues with high accuracy, but often fail to identify transitions between SOAP categories. We also report for the first time common trajectories in medical dialogue structures that provide valuable insights for designing `differential diagnosis' systems. Finally, we extensively study the impact of intent filtering for medical dialogue summarization and observe a significant boost in performance. We make the codes and data, including annotation guidelines, publicly available at https://github.com/DATEXIS/medical-intent-classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。