测试大模型对时间信息的鲁棒性,发现其在时间推理上表现差,改进后准确率提升55%。
A Study into Investigating Temporal Robustness of LLMs
- 设计8种时间敏感测试,评估大模型零样本下处理时间信息的能力
- 发现模型对时间表述变化和粒度差异极度敏感,性能下降明显
- 测试方法可实时评估用户提问的时间鲁棒性,适合改进问答系统
大型语言模型(LLMs)包含大量事实性世界知识,但在时间类问题和历史知识问答上表现受限,因其难以理解时间范围与方向,或完全忽略时间维度。本研究旨在精确衡量LLMs在基于时间信息的问答任务中,对时间推理和时间事实知识的鲁棒性。我们针对六种主流大模型,在零样本设置下设计了八项时间敏感的鲁棒性测试。结果表明,模型普遍缺乏时间鲁棒性,尤其在时间表述重构和不同时间粒度引用方面表现脆弱。我们展示如何利用其中部分测试自动实时判断模型对用户提问的时间鲁棒性。最后,基于研究发现,将时间问答性能最高提升55%。
原文摘要 · Abstract (English)
Large Language Models (LLMs) encapsulate a surprising amount of factual world knowledge. However, their performance on temporal questions and historical knowledge is limited because they often cannot understand temporal scope and orientation or neglect the temporal aspect altogether. In this study, we aim to measure precisely how robust LLMs are for question answering based on their ability to process temporal information and perform tasks requiring temporal reasoning and temporal factual knowledge. Specifically, we design eight time-sensitive robustness tests for factual information to check the sensitivity of six popular LLMs in the zero-shot setting. Overall, we find LLMs lacking temporal robustness, especially to temporal reformulations and the use of different granularities of temporal references. We show how a selection of these eight tests can be used automatically to judge a model's temporal robustness for user questions on the fly. Finally, we apply the findings of this study to improve the temporal QA performance by up to 55 percent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。