用图结构增强对话医学用药推荐,提升准确性和真实性
GAP: Graph-Assisted Prompts for Dialogue-based Medication Recommendation
- 构建患者中心的医疗概念图,捕捉对话中的细粒度信息
- 结合外部知识图谱生成多源查询,减少非事实性回答
- 在动态诊断访谈场景中表现优异,适合医疗对话系统研究者
用药推荐已成为医疗领域的重要任务,尤其在评估医疗对话系统(MDS)的准确性与安全性方面。与基于电子健康记录(EHR)的推荐不同,基于对话的用药推荐需关注患者与医生之间的交互细节,这些信息在EHR中可能缺失。尽管大语言模型(LLM)在医疗对话领域取得进展,能理解患者意图并提供包括用药建议在内的医学建议,但仍存在挑战:在多轮对话中,LLM可能忽略细粒度医疗信息或跨轮次的关联,且缺乏领域知识时易生成非事实性回应,这在医疗场景中尤为危险。为此,我们提出图辅助提示框架(GAP),从对话中提取医疗概念及其状态,构建显式的患者中心图,以描述被忽视但重要的信息。进一步结合外部医学知识图谱,GAP可生成丰富查询与提示,从而从多源检索信息,降低非事实性响应。我们在基于对话的用药推荐数据集上评估了GAP,并探索其在更复杂动态诊断访谈场景中的潜力。大量实验表明,GAP在性能上优于强基线方法。
原文摘要 · Abstract (English)
Medication recommendations have become an important task in the healthcare domain, especially in measuring the accuracy and safety of medical dialogue systems (MDS). Different from the recommendation task based on electronic health records (EHRs), dialogue-based medication recommendations require research on the interaction details between patients and doctors, which is crucial but may not exist in EHRs. Recent advancements in large language models (LLM) have extended the medical dialogue domain. These LLMs can interpret patients' intent and provide medical suggestions including medication recommendations, but some challenges are still worth attention. During a multi-turn dialogue, LLMs may ignore the fine-grained medical information or connections across the dialogue turns, which is vital for providing accurate suggestions. Besides, LLMs may generate non-factual responses when there is a lack of domain-specific knowledge, which is more risky in the medical domain. To address these challenges, we propose a \textbf{G}raph-\textbf{A}ssisted \textbf{P}rompts (\textbf{GAP}) framework for dialogue-based medication recommendation. It extracts medical concepts and corresponding states from dialogue to construct an explicitly patient-centric graph, which can describe the neglected but important information. Further, combined with external medical knowledge graphs, GAP can generate abundant queries and prompts, thus retrieving information from multiple sources to reduce the non-factual responses. We evaluate GAP on a dialogue-based medication recommendation dataset and further explore its potential in a more difficult scenario, dynamically diagnostic interviewing. Extensive experiments demonstrate its competitive performance when compared with strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。