arXiv:2501.07482cs.CLcs.AI2025-01被引 8

构建全球事件基准,评估大模型对重大时事的记忆能力

TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time

  • 基于维基百科结构化数据构建超2.3万条问答,覆盖10年全球事件
  • 发现模型在不同地区事实召回率差异显著,与各国人类发展指数相关性超0.7
  • 低资源语言表现明显下降,凸显多语言训练平衡的重要性

随着知识环境演变和大语言模型广泛应用,及时更新模型知识变得愈发重要。现有基准主要评估通用事实召回能力,却较少关注模型随时间推移或跨区域的知识保持情况。为此,我们提出时序事件基准(TiEBe),包含超过2.3万条问答对,聚焦过去10余年中全球及区域重大事件,覆盖23个地区、13种语言。该基准利用维基百科的结构化回溯数据识别重要事件,并据此构建评估体系,检验模型对全球与区域发展的理解,其依据为超越维基百科本身的客观事实证据。结果表明,模型在不同地区的事实召回存在显著地理差异,且其性能与各国人类发展指数(HDI)等社会经济指标的皮尔逊相关系数超过0.7。此外,以事件发生地母语提问,揭示出低资源语言下模型表现大幅下滑,凸显当前训练数据在全球分布上的不平衡。

原文摘要 · Abstract (English)

As the knowledge landscape evolves and large language models (LLMs) become increasingly widespread, there is a growing need to keep these models updated with current events. While existing benchmarks assess general factual recall, few studies explore how LLMs retain knowledge over time or across different regions. To address these gaps, we present the Timely Events Benchmark (TiEBe), a dataset of over 23,000 question-answer pairs centered on notable global and regional events, spanning more than 10 years of events, 23 regions, and 13 languages. TiEBe leverages structured retrospective data from Wikipedia to identify notable events through time. These events are then used to construct a benchmark to evaluate LLMs' understanding of global and regional developments, grounded in factual evidence beyond Wikipedia itself. Our results reveal significant geographic disparities in factual recall, emphasizing the need for more balanced global representation in LLM training. We also observe a Pearson correlation of more than 0.7 between models' performance in TiEBe and various countries' socioeconomic indicators, such as HDI. In addition, we examine the impact of language on factual recall by posing questions in the native language of the region where each event occurred, uncovering substantial performance gaps for low-resource languages.

大模型评估知识记忆多语言时序事件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。