arXiv:2510.19172cs.CLcs.AI2025-10被引 2

测试大模型对随时间变化的事实认知能力,发现性能下降超31%。

When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA

  • 基于真实时间数据构建动态知识评测集
  • 12个模型在动态问题上最高降31%准确率
  • 适合评估模型时序知识更新能力的研究者

大语言模型常无法处理训练数据中随时间演变的事实冲突。现有研究多基于维基数据等结构化知识库,聚焦热门实体且缺乏动态性,难以公平评估不同知识截止日期的模型。我们提出evolveQA,一个专为评估时序演化知识设计的基准,源自三个真实时间标注语料:AWS更新、Azure变更和世卫组织疫情报告。该框架识别自然发生的知识演进,生成与各模型知识截止日期匹配的黄金答案问题。对12个开源与闭源模型在三种知识探测格式上的评估显示,相比静态知识问题,其在evolveQA上性能最高下降31%。

原文摘要 · Abstract (English)

LLMs often fail to handle temporal knowledge conflicts--contradictions arising when facts evolve over time within their training data. Existing studies evaluate this phenomenon through benchmarks built on structured knowledge bases like Wikidata, but they focus on widely-covered, easily-memorized popular entities and lack the dynamic structure needed to fairly evaluate LLMs with different knowledge cut-off dates. We introduce evolveQA, a benchmark specifically designed to evaluate LLMs on temporally evolving knowledge, constructed from 3 real-world, time-stamped corpora: AWS updates, Azure changes, and WHO disease outbreak reports. Our framework identifies naturally occurring knowledge evolution and generates questions with gold answers tailored to different LLM knowledge cut-off dates. Through extensive evaluation of 12 open and closed-source LLMs across 3 knowledge probing formats, we demonstrate significant performance drops of up to 31% on evolveQA compared to static knowledge questions.

知识演化大模型评测时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。