用大模型生成假想病例,提升罕见药推荐准确率。
Improving Rare Medication Recommendation with Counterfactual Data Augmentation and Large Language Models

- 用大模型生成反事实医疗数据,缓解罕见药数据稀缺问题。
- 在罕见药推荐上性能比最强基线高出30.9%。
- 适合需要精准处理罕见病用药的临床AI研究者。
基于人工智能的药物推荐系统因其提升患者安全与治疗效果的潜力而备受关注。尽管准确推荐罕见处方药(rare-meds)具有重要临床意义,但现有方法在该任务上的预测性能显著偏低。我们归因于两个固有局限:(a) 罕见药数据本身稀缺;(b) 对共推荐药物关系考虑不足。为此,我们提出GenRxR框架,基于大语言模型(LLMs)构建。GenRxR利用LLM的医学知识与临床推理能力生成反事实医疗数据,缓解罕见药数据稀缺问题;同时将LLM融入推荐流程,建模共推荐药物间的关联。为进一步增强临床推理,引入指令微调步骤,使LLM能力与推荐任务对齐,更好处理包含罕见药的临床场景。实验表明,GenRxR在多数情况下优于14个基线(包括5个基于LLM的方法),在罕见药推荐上性能较最强基线最高提升30.9%。
原文摘要 · Abstract (English)
AI-based medication recommendation systems have attracted substantial attention due to their potential to enhance patient safety and therapeutic outcomes. Despite the clinical importance of accurately recommending rarely prescribed medications (rare-meds), we observe that most existing methods show significantly lower predictive performance for rare-meds. We attribute this issue to two intrinsic limitations: (a) the inherent scarcity of data for rare-meds and (b) limited consideration of co-recommended medications. To address these limitations, we propose GenRxR, a novel framework based on large language models (LLMs). GenRxR leverages the medical knowledge and clinical reasoning capability of LLMs to generate counterfactual medical data, mitigating the data scarcity issue for rare-meds. It also integrates an LLM into the medication recommendation process to model relationships among co-recommended medications. To further enhance the clinical reasoning, we introduce an instruction tuning step that aligns the LLM's capability with the recommendation task, enabling better handling of clinical context, including rare-meds cases. In our experiments, we show that GenRxR outperforms 14 (including 5 LLM-based) baselines in most cases. Specifically, it achieves up to 30.9% higher predictive performance for rare-meds than the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。