arXiv:2505.10798cs.CL2025-05被引 1

首个基于完整书籍的大型关系抽取数据集,助力社会科学研究。

Relation Extraction Across Entire Books to Reconstruct Community Networks: The AffilKG Datasets

  • 用整本书扫描配大标签知识图谱,构建人物-组织隶属关系网络
  • 六大数据集显示模型表现差异显著,验证了提取误差对分析的影响
  • 适合社会学、历史学研究者评估知识图谱在真实场景中的可靠性

当知识图谱(KG)从文本中自动提取时,其准确性是否足以支持下游分析?当前标注数据集无法回答此问题,因其知识图谱高度割裂、规模过小或过于复杂。为此,我们推出AffilKG(https://doi.org/10.5281/zenodo.15427977),这是首个将完整书籍扫描与大规模标注知识图谱配对的数据集集合。六个数据集均包含隶属关系图,即刻画人物与组织间成员关系的简单知识图谱,适用于移民、社群互动等社会现象研究;其中三个数据集还包含涵盖多种关系类型的扩展知识图谱。初步实验显示模型在不同数据集上表现差异显著,凸显AffilKG可实现两项关键进展:(1) 衡量关系抽取错误如何传播至图级分析(如社区结构);(2) 验证知识图谱提取方法在真实社会科学研究中的有效性。

原文摘要 · Abstract (English)

When knowledge graphs (KGs) are automatically extracted from text, are they accurate enough for downstream analysis? Unfortunately, current annotated datasets can not be used to evaluate this question, since their KGs are highly disconnected, too small, or overly complex. To address this gap, we introduce AffilKG (https://doi.org/10.5281/zenodo.15427977), which is a collection of six datasets that are the first to pair complete book scans with large, labeled knowledge graphs. Each dataset features affiliation graphs, which are simple KGs that capture Member relationships between Person and Organization entities -- useful in studies of migration, community interactions, and other social phenomena. In addition, three datasets include expanded KGs with a wider variety of relation types. Our preliminary experiments demonstrate significant variability in model performance across datasets, underscoring AffilKG's ability to enable two critical advances: (1) benchmarking how extraction errors propagate to graph-level analyses (e.g., community structure), and (2) validating KG extraction methods for real-world social science research.

知识图谱关系抽取社会计算数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。