arXiv:2410.21155cs.CL2024-10EMNLP被引 42

构建首个覆盖全篇科学文献的实体关系数据集,助力科研知识结构化。

SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific Documents

  • 基于106篇论文全文本标注,涵盖24000+实体与12000+关系。
  • 引入细粒度关系标签,捕捉方法、数据集、任务间的复杂交互。
  • 提供分布外测试集,推动更真实可靠的模型评估。

科学信息抽取(SciIE)对将学术文章中的非结构化知识转化为结构化数据至关重要。现有数据集多仅标注摘要等局部内容,导致上下文中的实体和关系缺失。本文发布首个面向科学文献中数据集、方法、任务相关实体与关系的抽取数据集。该数据集包含106篇人工标注的全文科学论文,涵盖超过24,000个实体和12,000条关系。为捕捉全文中实体间复杂的使用与交互关系,采用细粒度关系标签体系。同时提供分布外测试集,实现更贴近实际的评估。我们进行了全面实验,包括主流监督模型及自研大模型基线,揭示了该数据集带来的挑战,激励开发创新模型以推动SciIE领域发展。

原文摘要 · Abstract (English)

Scientific information extraction (SciIE) is critical for converting unstructured knowledge from scholarly articles into structured data (entities and relations). Several datasets have been proposed for training and validating SciIE models. However, due to the high complexity and cost of annotating scientific texts, those datasets restrict their annotations to specific parts of paper, such as abstracts, resulting in the loss of diverse entity mentions and relations in context. In this paper, we release a new entity and relation extraction dataset for entities related to datasets, methods, and tasks in scientific articles. Our dataset contains 106 manually annotated full-text scientific publications with over 24k entities and 12k relations. To capture the intricate use and interactions among entities in full texts, our dataset contains a fine-grained tag set for relations. Additionally, we provide an out-of-distribution test set to offer a more realistic evaluation. We conduct comprehensive experiments, including state-of-the-art supervised models and our proposed LLM-based baselines, and highlight the challenges presented by our dataset, encouraging the development of innovative models to further the field of SciIE.

信息抽取科学文献实体关系数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。