构建法律与人文学科引文数据集,提升脚注中引用信息的提取准确率
Digging Up Citations: FOSSIL, a Dataset and Workflow for Reference Extraction in Law and the Humanities

- 针对脚注引用混乱问题,设计多语言标注数据集和协作标注工具
- 专用管道使引文提取微平均F1从0.36提升至0.72,召回率显著改善
- 适合研究数字人文、法律文献挖掘及跨语言引文分析的学者使用
引文提取工具主要针对自然科学中结构化的文末参考文献,但法律与人文学科的引用多以脚注形式存在,其中书目信息与评论、交叉引用混杂,且语言和格式差异大。为解决高质量标注资源稀缺问题,本文提出FOSSIL(基于脚注的开放学术科学实例标签)数据集,包含96篇经标注的学术文章,共超过7,600个嵌入脚注的引用。同时提供PDF-TEI编辑器(协作式网页标注工具)、七标注员标准化工作流程文档,以及针对脚注引用的Grobid专项模型。在端到端评估中,该专用流水线将提取质量近乎翻倍,微平均F1从0.36提升至0.72,主要得益于召回率提高;但仍存在跨引用和混合内容脚注的改进空间。本摘要展示的是进行中的工作,引文分段、解析与交叉引用识别的标注仍在持续。
原文摘要 · Abstract (English)
Citation extraction tools are designed for the structured end-of-document bibliographies of the natural sciences, but law and humanities scholarship cites references primarily in footnotes, where bibliographic data is interleaved with commentary and cross-references and varies widely across languages and styles. To address the scarcity of suitable gold-standard resources, we present FOSSIL (Footnote-based Open-access SSH Scientific Instance Labels), an openly licensed multilingual dataset of 96 annotated scholarly articles containing over 7,600 footnote-embedded references, together with PDF-TEI Editor (a collaborative web annotation tool), a documented seven-annotator workflow, and a Grobid specialization for footnote-based citations. In end-to-end evaluation, the specialized pipeline nearly doubles extraction quality over default Grobid (micro-F1 from 0.36 to 0.72), driven largely by improved recall, while showing that substantial headroom remains for cross-references and mixed-content footnotes. This extended abstract presents work in progress; annotations of citations segmentation and parsing, and cross-reference resolution are ongoing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。