构建1000万+历史事件多语言数据集,提升大模型时序推理能力评测
HistoryBankQA: Multilingual Temporal Question Answering on Historical Events
- 从维基百科提取1000万+历史事件,覆盖10种语言
- 涵盖6类时序问答任务,验证大模型在跨语言时序理解上的表现
- 开源数据集与代码,助力多语言历史事件理解研究
对历史事件的时序推理是事件抽取、历史实体链接、时序问答、时间线摘要、时序事件聚类和时序自然语言推理等NLP任务的关键能力。然而,现有针对大语言模型(LLMs)时序推理能力的基准测试仍十分有限。现有时序推理数据集规模小、缺乏多语言覆盖,且更关注当代事件。为此,我们提出了HistoryBank,一个基于维基百科时间线页面和文章信息框提取的多语言历史事件数据库,包含超过1000万条历史事件,覆盖10种语言。该数据库在历史深度和语言广度上均达到前所未有的覆盖范围。此外,我们构建了一个全面的跨语言时序问答基准,涵盖6类时序问答推理任务,并评估了多种主流语言模型(LLaMA-3-8B、Mistral-7B、Gemma-2-9b、Qwen3-8B、GPT4o)在这些任务上的表现。结果显示,GPT4o在所有答案类型和语言中表现最佳;Gemma-2在小型模型中表现最优。本工作旨在为推进多语言、时序感知的历史事件自然语言理解提供综合性资源。为促进后续研究,我们将在论文接受后公开代码与数据集。
原文摘要 · Abstract (English)
Temporal reasoning about historical events is a critical skill for NLP tasks like event extraction, historical entity linking, temporal question answering, timeline summarization, temporal event clustering and temporal natural language inference. Yet efforts on benchmarking temporal reasoning capabilities of large language models (LLMs) are rather limited. Existing temporal reasoning datasets are limited in scale, lack multilingual coverage and focus more on contemporary events. To address these limitations, we present HistoryBank, a multilingual database of 10M+ historical events extracted from Wikipedia timeline pages and article infoboxes. Our database provides unprecedented coverage in both historical depth and linguistic breadth with 10 languages. Additionally, we construct a comprehensive question answering benchmark for temporal reasoning across all languages. This benchmark covers a diverse set of 6 temporal QA reasoning tasks, and we evaluate a suite of popular language models (LLaMA-3-8B, Mistral-7B, Gemma-2-9b, Qwen3-8B, GPT4o) to assess their performance on these tasks. As expected GPT4o performs best across all answer types and languages; Gemma-2 outperforms the other small language models. Our work aims to provide a comprehensive resource for advancing multilingual and temporally-aware natural language understanding of historical events. To facilitate further research, we will make our code and datasets publicly available upon acceptance of this paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。