对比两个数据集发现,LLM处理时间事实冲突的效果依赖数据设计和模型大小。
Temporal Fact Conflicts in LLMs: Reproducibility Insights from Unifying DYNAMICQA and MULAN
- 统一两个数据集评估标准,重现实验以分析结论分歧根源。
- 70%以上实验显示模型规模越大,更新时间事实能力越强。
- 数据集构造方式影响结果,自然语言上下文更贴近真实场景。
大型语言模型(LLMs)常因训练数据过时或信息演变而出现时间事实冲突。两项近期研究及其数据集得出了相反结论:DYNAMICQA认为外部上下文难以改变模型输出分布,时间事实更具抗性;而MULAN则发现外部上下文能频繁更新记忆中的时间事实。本文通过复现两份研究的实验,并在对方数据集上重新测试,探究分歧来源。为实现可比性,我们对两个数据集进行标准化处理。特别地,在复现DYNAMICQA时,使用大模型生成真实自然语言上下文,替代MULAN原有的程序化陈述。分析表明,结果高度依赖数据集设计:MULAN的发现可在两种框架下通用,但将MULAN评价方法应用于DYNAMICQA时结果不一致。此外,原研究仅针对7B模型,本文扩展至多尺寸模型,揭示模型规模对时间事实编码与更新的关键影响。结果强调数据集设计、评估指标和模型规模如何共同塑造LLM在时间知识冲突下的表现。
原文摘要 · Abstract (English)
Large Language Models (LLMs) often struggle with temporal fact conflicts due to outdated or evolving information in their training data. Two recent studies with accompanying datasets report opposite conclusions on whether external context can effectively resolve such conflicts. DYNAMICQA evaluates how effective external context is in shifting the model's output distribution, finding that temporal facts are more resistant to change. In contrast, MULAN examines how often external context changes memorised facts, concluding that temporal facts are easier to update. In this reproducibility paper, we first reproduce experiments from both benchmarks. We then reproduce the experiments of each study on the dataset of the other to investigate the source of their disagreement. To enable direct comparison of findings, we standardise both datasets to align with the evaluation settings of each study. Importantly, using an LLM, we synthetically generate realistic natural language contexts to replace MULAN's programmatically constructed statements when reproducing the findings of DYNAMICQA. Our analysis reveals strong dataset dependence: MULAN's findings generalise under both methodological frameworks, whereas applying MULAN's evaluation to DYNAMICQA yields mixed outcomes. Finally, while the original studies only considered 7B LLMs, we reproduce these experiments across LLMs of varying sizes, revealing how model size influences the encoding and updating of temporal facts. Our results highlight how dataset design, evaluation metrics, and model size shape LLM behaviour in the presence of temporal knowledge conflicts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。