arXiv:2603.25640cs.DLcs.CL2026-03

构建首个公开可复现的引文解析基准,助力学术自动化

RenoBench: A Citation Parsing Benchmark

  • 从16万条标注引文出发,经自动化验证生成1万条多语言跨平台数据
  • 评估多种解析系统,大模型微调后在各领域精度超90%
  • 适合研究学术基础设施、自然语言处理与科学计量的开发者

准确解析引文是实现机器可读学术基础设施的关键。然而,现有评估方法常缺乏通用性、依赖合成数据或未公开。我们提出RenoBench,一个基于四大出版生态(SciELO、Redalyc、Public Knowledge Project、Open Research Europe)PDF文档的公开引文解析基准。从16.1万条标注引文出发,通过自动化验证与基于特征的采样,构建了涵盖多语言、多出版类型、多平台的1万条引文数据集。我们评估了多种引文解析系统,并报告了各领域的精确率与召回率。结果表明,经过微调的语言模型表现优异,尤其在复杂格式下仍保持高精度。RenoBench支持可复现、标准化的评估,为自动引文解析与元科学研究提供基础。

原文摘要 · Abstract (English)

Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, existing evaluation techniques are often not generalizable, based on synthetic data, or not publicly available. We introduce RenoBench, a public domain benchmark for citation parsing, sourced from PDFs released on four publishing ecosystems: SciELO, Redalyc, the Public Knowledge Project, and Open Research Europe. Starting from 161,000 annotated citations, we apply automated validation and feature-based sampling to produce a dataset of 10,000 citations spanning multiple languages, publication types, and platforms. We then evaluate a variety of citation parsing systems and report field-level precision and recall. Our results show strong performance from language models, particularly when fine-tuned. RenoBench enables reproducible, standardized evaluation of citation parsing systems, and provides a foundation for advancing automated citation parsing and metascientific research.

引文解析数据集NLP学术智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。