arXiv:2509.24283cs.DLcs.CL2025-09综述

构建科学文献引文预测新基准,涵盖发现、补全与定位三任务。

Overview of SCIDOCA 2025 Shared Task on Citation Prediction, Discovery, and Placement

  • 分三阶段评估引文识别:发现相关参考文献、补全掩码引文、判断引文句归属。
  • 基于S2ORC构建超6万段标注数据集,测试集含1000段独立论文段落。
  • 七队参赛,三队提交结果,公开数据与评测工具促科研进步。

本文介绍SCIDOCA 2025共享任务,聚焦科学文档中的引文发现与预测。任务分为三个子任务:(1) 引文发现——为给定段落识别相关参考文献;(2) 掩码引文预测——从候选中选出正确引文填补空缺;(3) 引文句预测——确定每句引用的准确来源。我们基于语义学者开放研究语料库(S2ORC)构建大规模数据集,包含超过60,000个标注段落及精选参考文献集。测试集由1,000段来自不同论文的段落组成,每段均配有真实引文与干扰项。共有七支队伍注册,其中三支提交结果。报告各子任务性能指标,并分析系统有效性。该共享任务为引文建模提供新基准,推动科学文档理解研究。数据集与任务材料已公开于 https://github.com/daotuanan/scidoca2025-shared-task。

原文摘要 · Abstract (English)

We present an overview of the SCIDOCA 2025 Shared Task, which focuses on citation discovery and prediction in scientific documents. The task is divided into three subtasks: (1) Citation Discovery, where systems must identify relevant references for a given paragraph; (2) Masked Citation Prediction, which requires selecting the correct citation for masked citation slots; and (3) Citation Sentence Prediction, where systems must determine the correct reference for each cited sentence. We release a large-scale dataset constructed from the Semantic Scholar Open Research Corpus (S2ORC), containing over 60,000 annotated paragraphs and a curated reference set. The test set consists of 1,000 paragraphs from distinct papers, each annotated with ground-truth citations and distractor candidates. A total of seven teams registered, with three submitting results. We report performance metrics across all subtasks and analyze the effectiveness of submitted systems. This shared task provides a new benchmark for evaluating citation modeling and encourages future research in scientific document understanding. The dataset and task materials are publicly available at https://github.com/daotuanan/scidoca2025-shared-task.

引文预测科学文本共享任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。