arXiv:2412.13582cs.CL2024-12被引 15

构建动态知识评估数据集,测试大模型如何跟上知识变化。

EvoWiki: Evaluating LLMs on Evolving Knowledge

  • 用稳定、演变、未知三类信息构建可自动更新的动态数据集。
  • 现有模型对新知识适应差,常给出过时或错误答案。
  • 检索增强与持续学习结合能更好应对知识演化,适合研究者参考。

大语言模型的知识利用能力至关重要,理解其对不断演化的知识的适应性对有效部署至关重要。然而,现有基准大多为静态,无法捕捉模型与知识的动态特性,导致准确性下降和数据污染等问题。本文提出EvoWiki,一个通过将信息划分为稳定、演化和未探索状态来反映知识演化的动态数据集,具备完全自动化更新能力,可精准评估持续变化的知识及新发布的大模型。我们通过检索增强生成(RAG)和持续学习(CL)实验,评估模型对演化知识的适应能力。结果表明,当前模型在处理演化知识时常表现不佳,频繁提供过时或错误回答。此外,数据集揭示了RAG与CL之间的协同效应,展示了二者结合在应对知识演化方面的潜力。EvoWiki为推进大模型知识演化能力的研究提供了可靠基准。

原文摘要 · Abstract (English)

Knowledge utilization is a critical aspect of LLMs, and understanding how they adapt to evolving knowledge is essential for their effective deployment. However, existing benchmarks are predominantly static, failing to capture the evolving nature of LLMs and knowledge, leading to inaccuracies and vulnerabilities such as contamination. In this paper, we introduce EvoWiki, an evolving dataset designed to reflect knowledge evolution by categorizing information into stable, evolved, and uncharted states. EvoWiki is fully auto-updatable, enabling precise evaluation of continuously changing knowledge and newly released LLMs. Through experiments with Retrieval-Augmented Generation (RAG) and Contunual Learning (CL), we evaluate how effectively LLMs adapt to evolving knowledge. Our results indicate that current models often struggle with evolved knowledge, frequently providing outdated or incorrect responses. Moreover, the dataset highlights a synergistic effect between RAG and CL, demonstrating their potential to better adapt to evolving knowledge. EvoWiki provides a robust benchmark for advancing future research on the knowledge evolution capabilities of large language models.

知识演化大模型评估动态数据集RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。