新术语基准测试揭示大模型对实时信息的严重滞后问题
NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates
- 构建自动化更新的基准NewTerm,支持年度动态评测新术语
- 实测显示新术语导致大模型性能下降超20%
- 揭示知识更新难以泛化至远期新词,为改进指明方向
尽管大语言模型在多项任务中表现优异,但因训练过程存在知识截止问题,难以处理实时信息(如新事实和新术语)。现有基准多聚焦过时内容且覆盖领域有限,难以实现动态更新,忽视了新术语的评估。为此,我们提出自适应基准NewTerm,设计高度自动化的构建方法,在极少人工干预下实现高质量基准建设,并支持实时更新。在多个大语言模型上的实验证明,新术语导致性能下降超过20%。尽管更新知识截止可覆盖部分新术语,但无法泛化至更远期的新词。我们进一步分析了哪些类型术语更难处理,揭示了模型应对新术语的局限性,为后续研究铺路。目前已构建NewTerm 2022与2023,将持续每年更新。基准与代码见https://github.com/hexuandeng/NewTerm。
原文摘要 · Abstract (English)
Despite their remarkable abilities in various tasks, large language models (LLMs) still struggle with real-time information (e.g., new facts and terms) due to the knowledge cutoff in their development process. However, existing benchmarks focus on outdated content and limited fields, facing difficulties in real-time updating and leaving new terms unexplored. To address this problem, we propose an adaptive benchmark, NewTerm, for real-time evaluation of new terms. We design a highly automated construction method to ensure high-quality benchmark construction with minimal human effort, allowing flexible updates for real-time information. Empirical results on various LLMs demonstrate over 20% performance reduction caused by new terms. Additionally, while updates to the knowledge cutoff of LLMs can cover some of the new terms, they are unable to generalize to more distant new terms. We also analyze which types of terms are more challenging and why LLMs struggle with new terms, paving the way for future research. Finally, we construct NewTerm 2022 and 2023 to evaluate the new terms updated each year and will continue updating annually. The benchmark and codes can be found at https://github.com/hexuandeng/NewTerm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。