arXiv:2605.30582cs.CL2026-05

用大模型自动追踪论文中数据集使用情况,提升科研透明度。

AI for Monitoring and Classifying Data Used in Research Literature

论文配图:AI for Monitoring and Classifying Data Used in Research Literature
图 1 · 摘自论文原文
  • 基于GLiNER的多任务框架,同时识别数据集、关系和使用上下文。
  • 通过合成数据与大模型重校验,解决标注数据少、引用模糊问题。
  • 适合关注科研可复现性与数据影响力的研究者使用。

尽管谷歌学术和Semantic Scholar等平台能追踪论文引用,但缺乏对研究文献中数据集使用的监控体系,导致数据使用情况不透明。这一空白影响科研透明性、可复现性及影响力评估,而现有进展受限于引用不规范、标注数据稀缺以及真实文本中数据集引用模糊等问题。传统NLP方法难以应对,促使转向更适应性强、语义丰富的模型。本文基于先前使用LLM进行数据提及检测和合成数据预训练的工作,提出一种可扩展的数据集监控新方法。引入基于GLiNER的多任务框架,联合执行数据集提及抽取、关系识别与使用上下文分类。针对标注稀缺问题,该流程结合合成数据生成与LLM-based重校验,过滤错误提及并确保标注一致性,从而提升训练管道的可靠性、覆盖范围与输出一致性。本工作推动了开源工具在科研文献数据使用监测中的发展,助力实现通用化、无约束的数据集引用追踪。

原文摘要 · Abstract (English)

While platforms like Google Scholar and Semantic Scholar track citations for academic papers, no comparable infrastructure exists for monitoring dataset usage in research literature, leaving the landscape of data use largely opaque. Addressing this gap is critical for transparency, reproducibility, and monitoring of impact, yet progress is hindered by inconsistent citation practices, scarce labeled data, and ambiguous references to datasets in the wild. Traditional NLP approaches struggle with these challenges, motivating the shift toward more adaptive, semantically rich models. Building on prior work using LLMs for data mention detection and synthetic data for bootstrapping training, this paper presents an updated methodology for scalable dataset monitoring. We introduce a multitask GLiNER-based framework that jointly performs dataset mention extraction, relation identification, and usage-context classification. To address label scarcity, the pipeline leverages synthetic data generation to produce training examples and LLM-based revalidation to filter incorrect mentions and enforce labeling consistency, together improving reliability, coverage, and output consistency across the training pipeline. This work advances the development of open-source tools for monitoring data use in research literature, contributing to the broader goal of generalizable, unconstrained dataset citation tracking.

数据追踪大模型科研可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。