arXiv:2605.13045cs.LGcs.CL2026-05

现有大模型缺乏对医学知识时效性的认知,导致判断过时或更新信息时出错。

Large Language Models Lack Temporal Awareness of Medical Knowledge

论文配图:Large Language Models Lack Temporal Awareness of Medical Knowledge
图 1 · 摘自论文原文
  • 构建首个基于动态指南的医学时效性评估基准TempoMed-Bench
  • 模型对最新医学知识准确率随时间线性下降,旧知识回忆准确率仅为新知识的25.37%-53.89%
  • 即使使用搜索工具也无法有效缓解时效认知偏差,适合医学AI研究者关注

当前评估大语言模型(LLMs)医学知识的方法多依赖静态测试集,而医学知识本身具有动态性,随新证据出现和治疗获批持续演进。因此,缺乏时间上下文的评估可能无法真实反映模型对特定时间节点医学知识的推理能力。多数医学数据为历史数据,要求模型不仅记住正确知识,还需知晓其适用时间。为此,我们构建了首个面向医学领域的时间感知评估基准TempoMed-Bench,通过演化指南知识评估模型时效意识。基于该基准的分析揭示:(1)模型在最新医学知识上的表现随时间呈渐进式线性下降,而非突变式知识截止,表明参数化医学知识未严格受知识截止限制;(2)模型在回忆历史过时知识时表现更差,准确率仅为最新推荐知识的25.37%-53.89%,暗示训练中存在潜在知识遗忘现象;(3)模型常表现出时间不一致行为,预测结果在相邻年份间波动无规律。此外,即便结合代理搜索工具,时效意识问题仍难以解决,性能下降3.15%-14.14%。本工作揭示了一个重要但被忽视的挑战,推动未来研究发展具备时间敏感医学知识编码能力的LLMs。

原文摘要 · Abstract (English)

The existing methods for evaluating the medical knowledge of Large Language Models (LLMs) are largely based on atemporal examination-style benchmarks, while in reality, medical knowledge is inherently dynamic and continuously evolves as new evidence emerges and treatments are approved. Consequently, evaluating medical knowledge without a temporal context may provide an incomplete assessment of whether LLMs can accurately reason about time-specific medical knowledge. Moreover, most medical data are historical, requiring the models not only to recall the correct knowledge, but also to know when that knowledge is correct. To bridge the gap, we built TempoMed-Bench, the first-of-its-kind benchmark for evaluating the temporal awareness of the LLMs in the medical domain through evolving guideline knowledge. Based on the TempoMed-Bench, our evaluation analysis first reveals that LLMs lack temporal awareness in medical knowledge through the key findings: (1) model performance on up-to-date medical knowledge exhibits a gradual linear decline over time rather than a sharp knowledge-cutoff behavior, suggesting that parametric medical knowledge is not strictly bounded by knowledge cutoffs; (2) LLMs consistently struggle more with recalling outdated historical medical knowledge than with up-to-date recommendations: accuracy of historical knowledge is only 25.37%-53.89% of up-to-date knowledge, indicating potential knowledge forgetting effects during training; and (3) LLMs often exhibit temporally inconsistent behaviors, where predictions fluctuate irregularly across neighboring years. We also show that the temporal awareness problem is a challenge that cannot be easily solved when integrated with agentic search tools (-3.15%-14.14%). This work highlights an important yet underexplored challenge and motivates future research on developing LLMs that can better encode time-specific medical knowledge.

医学AI时效性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。