arXiv:2601.05099cs.DLcs.CL2026-01中稿 · the 25th ACM/IEEE …被引 4

从论文引用上下文挖掘真实研究用的数据集,提升发现准确率。

Multi-Disciplinary Dataset Discovery from Citation-Verified Literature Contexts

  • 通过分析论文引用语境识别数据集,不依赖低质量元数据。
  • 在8个计算机科学问题上平均召回率达47.47%,最高达81.82%。
  • 可发现未被文献记录的高价值新数据集,适合科研人员使用。

确定研究问题适用的数据集仍具挑战性,因现有搜索工具高度依赖元数据质量和关键词匹配,常无法捕捉科研意图。本文提出一种基于文献的框架,从科学论文的引用上下文中发现数据集,实现基于实际研究使用的检索。方法结合大规模引用上下文抽取、基于模式的大型语言模型数据集识别以及保留来源信息的实体消歧。在8个由综述导出的计算机科学查询上评估,系统召回率显著高于Google Dataset Search和DataCite Commons,平均为47.47%,最高达81.82%。除恢复标准数据集外,还发现了未在综述中记录的额外数据集。专家在五个一级学科领域评估认为,部分新增数据集具有高实用价值,甚至对特定主题属新颖发现。结果表明,引用上下文挖掘是数据集发现的有效通用范式,尤其适用于元数据不足或不可靠的情形。为支持可复现性与扩展,代码、评估数据集及结果已开源(https://github.com/Fireblossom/citation-context-dataset-discovery)。

原文摘要 · Abstract (English)

Identifying suitable datasets for a research question remains challenging because existing dataset search engines rely heavily on metadata quality and keyword overlap, which often fail to capture the semantic intent of scientific investigation. We introduce a literature-driven framework that discovers datasets from citation contexts in scientific papers, enabling retrieval grounded in actual research use rather than metadata availability. Our approach combines large-scale citation-context extraction, schema-guided dataset recognition with Large Language Models, and provenance-preserving entity resolution. We evaluate the system on eight survey-derived computer science queries and find that it achieves substantially higher recall than Google Dataset Search and DataCite Commons, with normalized recall ranging from an average of 47.47% to a highest value of 81.82%. Beyond recovering gold-standard datasets, the method also surfaces additional datasets not documented in the surveys. Expert assessments across five top-level Fields of Science indicate that a substantial portion of the additional datasets are considered high utility, and some are regarded as novel for the specific topics chosen by the experts. These findings establish citation-context mining as an effective and generalizable paradigm for dataset discovery, particularly in settings where datasets lack sufficient or reliable metadata. To support reproducibility and future extensions, we release our code, evaluation datasets, and results on GitHub (https://github.com/Fireblossom/citation-context-dataset-discovery).

数据集发现引文挖掘LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。