检测并缓解大模型在医疗知识演进中的过时与冲突问题
Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
- 构建DriftMedQA基准,模拟临床指南演化评估模型时效性
- 7个主流模型在4290个场景中普遍存在采纳过时建议和矛盾指引
- 检索增强生成+偏好微调可显著提升模型一致性与可靠性
大型语言模型(LLMs)在医疗领域潜力巨大,但难以适应快速演进的医学知识,导致推荐过时或相互矛盾。本研究探究了LLMs对不断变化的临床指南的响应情况,重点关注概念漂移与内部不一致问题。我们开发了DriftMedQA基准以模拟指南演变,并评估了多个LLM在时间维度上的可靠性。在4,290个场景下对7个前沿模型的评估显示,它们普遍难以拒绝过时建议,且频繁支持冲突指导。此外,我们探索了两种缓解策略:检索增强生成(RAG)和通过直接偏好优化(DPO)进行偏好微调。尽管每种方法均提升了性能,但二者结合效果最佳,实现了最一致可靠的输出。研究强调需提升模型对时间性变化的鲁棒性,以保障其在临床实践中的可信应用。数据集已公开于https://huggingface.co/datasets/RDBH/DriftMed。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have great potential in the field of health care, yet they face great challenges in adapting to rapidly evolving medical knowledge. This can lead to outdated or contradictory treatment suggestions. This study investigated how LLMs respond to evolving clinical guidelines, focusing on concept drift and internal inconsistencies. We developed the DriftMedQA benchmark to simulate guideline evolution and assessed the temporal reliability of various LLMs. Our evaluation of seven state-of-the-art models across 4,290 scenarios demonstrated difficulties in rejecting outdated recommendations and frequently endorsing conflicting guidance. Additionally, we explored two mitigation strategies: Retrieval-Augmented Generation and preference fine-tuning via Direct Preference Optimization. While each method improved model performance, their combination led to the most consistent and reliable results. These findings underscore the need to improve LLM robustness to temporal shifts to ensure more dependable applications in clinical practice. The dataset is available at https://huggingface.co/datasets/RDBH/DriftMed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。