构建法语历史问答数据集,支持多跳推理与跨源分析。
HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)

- 从法国第三共和国议会辩论中构建多跳问题,融合多源史料。
- 含1782个问题,强调时间推理与稀疏证据整合的复杂性。
- 适合历史研究与NLP模型在专业领域评估的场景。
我们提出HistoriQA-ThirdRepublic:一个基于法国第三共和国(1870–1940)议会辩论与报纸的法语历史问答语料库。该数据集由历史学者协作设计,捕捉历史研究中的典型复杂推理模式,包括跨源信息融合、时间推理及稀疏证据整合。共包含1782个问题,强调异构历史文档间的多跳关联,可用于评估检索增强型与大语言模型在特定领域任务中的表现。论文详细描述了语料构建方法,涵盖资料选择与对齐、问题验证及元数据集成。尽管聚焦法语文本,其方法可扩展至其他语言与国家语料库。最后,展示了该数据集如何支持真实的多跳问答评估场景,弥合NLP基准测试与历史学术需求之间的差距。
原文摘要 · Abstract (English)
We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic. Designed in collaboration with a historian, the corpus captures complex reasoning patterns typical of historical inquiry, including cross-source synthesis, temporal reasoning, and the integration of sparse evidence. The dataset is made of 1782 questions and emphasizes multi-hop connections across heterogeneous historical documents, providing a resource for evaluating retrieval-augmented and large language model systems in domain-specific contexts. We describe the methodology for constructing the corpus, including the selection and alignment of sources, question validation, and metadata integration. While the dataset focuses on French historical documents, our methodology can be readily adapted to other languages and national corpora. Finally, we demonstrate how the corpus can support realistic evaluation scenarios for multi-hop question answering, bridging the gap between NLP benchmarks and the needs of historical scholarship.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。