提出新方法评估大模型随时间演化的知识,发现其记忆存在时间边界断裂问题。
ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains
- 基于分阶段时间跨度的提示法,挖掘模型对时序知识的非参数化记忆
- 模型在训练数据格式不同下,回忆时间知识的能力差异明显
- 适用于检测模型对科学、法律等动态知识的时效性,适合评测者使用
大语言模型(LLMs)已深刻影响多个领域,但对其时序知识的评估与保障仍具挑战。现有方法多依赖固定时间点视角,难以捕捉知识的动态演化。为此,我们构建了ChroKnowBench基准数据集,用于评估跨多个领域、具有时间依赖性和时态状态的累积知识。该数据集区分了随时间变化的知识(如个人经历、科学发现、法律修订)与恒定不变的知识(如数学真理、常识)。基于此,我们提出ChroKnowledge框架——一种基于采样的非参数化时序知识评估方法。实验发现:(1)模型提取时序知识的能力与其训练数据格式密切相关;(2)模型常在时间边界处出现记忆断层,而非完整召回。为此,我们设计ChroKnowPrompt,通过逐步遍历时间邻域进行深度提示,显著提升对开源与专有模型中对象的跨时期召回能力,验证了其通用性。然而,在动态数据集和非结构化格式下仍面临挑战。
原文摘要 · Abstract (English)
Large language models (LLMs) have brought significant changes to many aspects of our lives. However, assessing and ensuring their chronological knowledge remains challenging. Existing approaches fall short in addressing the temporal adaptability of knowledge, often relying on a fixed time-point view. To overcome this, we introduce ChroKnowBench, a benchmark dataset designed to evaluate chronologically accumulated knowledge across three key aspects: multiple domains, time dependency, temporal state. Our benchmark distinguishes between knowledge that evolves (e.g., personal history, scientific discoveries, amended laws) and knowledge that remain constant (e.g., mathematical truths, commonsense facts). Building on this benchmark, we present ChroKnowledge (Chronological Categorization of Knowledge), a novel sampling-based framework for evaluating LLMs' non-parametric chronological knowledge. Our evaluation led to the following observations: (1) The ability of eliciting temporal knowledge varies depending on the data format that model was trained on. (2) LLMs partially recall knowledge or show a cut-off at temporal boundaries rather than recalling all aspects of knowledge correctly. Thus, we apply our ChroKnowPrompt, an in-depth prompting to elicit chronological knowledge by traversing step-by-step through the surrounding time spans. We observe that it successfully recalls objects across both open-source and proprietary LLMs, demonstrating versatility, though it faces challenges with dynamic datasets and unstructured formats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。