首个评估大模型用药安全的基准测试,发现其对隐含风险识别能力差。
RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
- 构建模拟问诊场景与用药风险数据库,生成2443个高质量案例
- 6725种禁忌、28781种药物相互作用,模型对隐含风险识别率低
- 适合医疗AI安全研究者,助力提升临床决策系统可靠性
基于大语言模型(LLMs)的医疗系统在多项任务中取得显著进展,但因真实数据受限于隐私与可及性,其用药安全性研究仍不充分。为填补空白,本文提出一个模拟临床问诊并评估用药安全性的框架。通过生成嵌入用药风险的问诊对话,构建包含6,725种禁忌、28,781种药物相互作用和14,906对适应症-药物关系的专用数据库RxRisk DB。采用两阶段筛选策略确保临床真实性和专业质量,最终形成包含2,443个高质量咨询场景的基准测试集RxSafeBench。使用结构化多选题评估主流开源与专有模型在模拟患者情境下推荐安全药物的能力。结果表明,当前模型难以整合禁忌与相互作用知识,尤其在风险隐含而非明确时表现更差。研究揭示了大模型在医疗应用中的关键安全挑战,并为通过提示优化与任务特定调优提升可靠性提供依据。该基准首次实现对大模型用药安全的系统性评估,推动更安全可信的AI临床决策支持发展。
原文摘要 · Abstract (English)
Numerous medical systems powered by Large Language Models (LLMs) have achieved remarkable progress in diverse healthcare tasks. However, research on their medication safety remains limited due to the lack of real world datasets, constrained by privacy and accessibility issues. Moreover, evaluation of LLMs in realistic clinical consultation settings, particularly regarding medication safety, is still underexplored. To address these gaps, we propose a framework that simulates and evaluates clinical consultations to systematically assess the medication safety capabilities of LLMs. Within this framework, we generate inquiry diagnosis dialogues with embedded medication risks and construct a dedicated medication safety database, RxRisk DB, containing 6,725 contraindications, 28,781 drug interactions, and 14,906 indication-drug pairs. A two-stage filtering strategy ensures clinical realism and professional quality, resulting in the benchmark RxSafeBench with 2,443 high-quality consultation scenarios. We evaluate leading open-source and proprietary LLMs using structured multiple choice questions that test their ability to recommend safe medications under simulated patient contexts. Results show that current LLMs struggle to integrate contraindication and interaction knowledge, especially when risks are implied rather than explicit. Our findings highlight key challenges in ensuring medication safety in LLM-based systems and provide insights into improving reliability through better prompting and task-specific tuning. RxSafeBench offers the first comprehensive benchmark for evaluating medication safety in LLMs, advancing safer and more trustworthy AI-driven clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。