arXiv:2606.22645cs.IRcs.CY2026-06被引 1

构建可问答的罗马帝国知识图谱,统一文本与图谱检索

All Relations Lead to Rome: Automated Knowledge Graph Creation and Question Generation

论文配图:All Relations Lead to Rome: Automated Knowledge Graph Creation and Question Generation
图 1 · 摘自论文原文
  • 联合生成知识图谱、嵌入向量与事实问答对
  • 覆盖19000实体、8400问答对,支持精准溯源
  • 适合研究混合检索与语义引导系统的学者

大型语言模型显著提升了信息检索与问答能力,但现有数据集通常仅支持基于向量的非结构化文本检索,或基于知识图谱的推理,缺乏融合两种范式的统一表示。此外,当前评估基准很少提供与原始语料对齐的实体、关系及事实支撑的问答对。为此,我们提出统一框架 All Relations Lead to Rome (ARLtR),实现知识图谱、嵌入表示与事实问答对的自动化联合构建。该框架显式地基于抽取的实体、关系和文本证据生成内容。我们进一步将框架实例化为以罗马帝国为中心的历史数据集,包含超过19,000个实体、16,000段文本块和8,400个问答对(https://huggingface.co/datasets/FaynePro/all-relations-lead-to-rome)。通过紧密耦合符号图结构与密集检索表示,ARLtR为混合检索系统与语义引导方法的评估与开发提供了单一连贯资源。

原文摘要 · Abstract (English)

Large language models have substantially improved information retrieval and question answering; however, existing datasets generally support either vector-based retrieval over unstructured text or reasoning over knowledge graphs, without providing a unified representation that combines both paradigms. Moreover, current benchmarks rarely provide ground-truth entities, relations, and fact-grounded question-answer pairs aligned with the underlying corpus. To address this gap, we introduce All Relations Lead to Rome (ARLtR), a unified framework for automated knowledge graph construction and fact-grounded question-answer generation. ARLtR jointly constructs a knowledge graph, embeddings, and question-answer pairs that are explicitly grounded in extracted entities, relations, and supporting textual evidence. We further instantiate the framework as a historical dataset centered on the Roman Empire, comprising over 19,000 entities, 16,000 chunks, and 8,400 question-answer pairs (https://huggingface.co/datasets/FaynePro/all-relations-lead-to-rome). By tightly coupling symbolic graph representations with dense retrieval representations, ARLtR facilitates the evaluation and development of hybrid retrieval systems and semantic steering approaches within a single coherent resource.

知识图谱问答生成混合检索历史数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。