arXiv:2604.28061cs.DLcs.CL2026-04中稿 · following feedback…被引 1

用大模型测科研数据重用率,发现43%论文重用了数据。

Measuring research data reuse in scholarly publications using generative artificial intelligence: Open Science Indicator development and preliminary results

  • 基于大语言模型构建数据重用指标,自动识别论文中对研究数据的引用与使用。
  • 测得科研数据重用率达43%,高于传统文献计量方法结果。
  • 为评估开放科学实际影响提供可扩展的新工具,适合政策制定者和期刊参考。

众多元科学研究及其他倡议已开始监测开放科学实践的普及情况,但更需关注开放科学的‘下游’影响。PLOS与DataSeer开发了一种基于大语言模型(LLM)的新指标,用于衡量开放科学的重要效果之一:研究数据的重用。结果显示,数据重用率为43%,高于现有文献计量技术。研究表明,利用大语言模型与生成式人工智能可在大规模上有效测量数据重用。当前研究数据共享与重用的积极影响可能被低估。

原文摘要 · Abstract (English)

Numerous metascience studies and other initiatives have begun to monitor the prevalence of open science practices when it is more important to understand the 'downstream' effects or impacts of open science. PLOS and DataSeer have developed a new LLM-based indicator to measure an important effect of open science: the reuse of research data. Our results show a data reuse rate of 43%, which is higher than established bibliometric techniques. We show that data reuse can be measured at scale using LLMs and generative artificial intelligence. The positive effects of research data sharing and reuse may currently be underestimated.

数据重用大模型开放科学指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。