arXiv:2509.04304cs.CLcs.AI2025-09EMNLP被引 7

测试大模型对过时医学知识的记忆,发现普遍给出过时建议。

Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models

  • 构建两个基于系统综述的问答数据集,检测模型对过时知识的记忆。
  • 8个主流大模型均在新旧共识对比任务中错误坚持旧知识,准确率下降超40%。
  • 适合医疗AI安全评估、医学大模型开发者参考,关注知识时效性者必读。

大型语言模型(LLM)在医疗领域的应用潜力巨大,但其依赖静态训练数据的特性带来风险:当医学共识随研究更新时,模型可能仍记忆并输出过时知识,导致有害建议或临床推理失败。为此,我们基于系统综述构建了两个新问答数据集:MedRevQA(16,501对问答,涵盖通用生物医学知识)和MedChangeQA(512对问答,聚焦医学共识变迁)。对8个主流大模型的评估显示,所有模型均持续依赖过时知识。我们进一步分析了过时预训练数据与训练策略的影响,并提出缓解方向,为构建更及时、可靠的医疗AI系统奠定基础。

原文摘要 · Abstract (English)

The growing capabilities of Large Language Models (LLMs) show significant potential to enhance healthcare by assisting medical researchers and physicians. However, their reliance on static training data is a major risk when medical recommendations evolve with new research and developments. When LLMs memorize outdated medical knowledge, they can provide harmful advice or fail at clinical reasoning tasks. To investigate this problem, we introduce two novel question-answering (QA) datasets derived from systematic reviews: MedRevQA (16,501 QA pairs covering general biomedical knowledge) and MedChangeQA (a subset of 512 QA pairs where medical consensus has changed over time). Our evaluation of eight prominent LLMs on the datasets reveals consistent reliance on outdated knowledge across all models. We additionally analyze the influence of obsolete pre-training data and training strategies to explain this phenomenon and propose future directions for mitigation, laying the groundwork for developing more current and reliable medical AI systems.

医疗AI知识时效大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。