用NLP分析论文全文,量化开放科研平台的实际影响。
Assessing the impact of Open Research Information Infrastructures using NLP driven full-text Scientometrics: A case study of the LXCat open-access platform
- 通过NLP提取论文中数据使用模式,揭示平台真实使用情况。
- 发现十年间研究者对LXCat平台的数据依赖度持续上升。
- 方法可迁移至其他科研平台,适合评估开放基础设施价值。
开放科研信息(ORI)在科学知识的生成、传播、验证与重用中起核心作用。传统评估多依赖引用指标,本文提出一种基于自然语言处理(NLP)的全文科学计量框架,以量化非引用层面的影响力,以低温等离子体(LTP)领域的开源平台LXCat为案例。通过对引用三篇基础LXCat文献的全文论文进行分析,构建整合化学实体识别、数据集与求解器提及提取、机构地理映射及主题建模的全流程方法,揭示数据使用中的细粒度模式,包括研究偏好、数据实践差异、对特定数据库的依赖程度、数据重用方式演变及科学工作流耦合关系。该方法具有领域无关性,可迁移至其他开放科研基础设施场景,为数据驱动地评估开源平台对研究方向的影响提供新工具。本框架具备可扩展性,可用于支持基础设施的证据评估、设计优化与政策制定。
原文摘要 · Abstract (English)
Open research information (ORI) play a central role in shaping how scientific knowledge is produced, disseminated, validated, and reused across the research lifecycle. While the visibility of such ORI infrastructures is often assessed through citation-based metrics, in this study, we present a full-text, natural language processing (NLP) driven scientometric framework to systematically quantify the impact of ORI infrastructures beyond citation counts, using the LXCat platform for low temperature plasma (LTP) research as a representative case study. The modeling of LTPs and interpretation of LTP experiments rely heavily on accurate data, much of which is hosted on LXCat, a community-driven, open-access platform central to the LTP research ecosystem. To investigate the scholarly impact of the LXCat platform over the past decade, we analyzed a curated corpus of full-text research articles citing three foundational LXCat publications. We present a comprehensive pipeline that integrates chemical entity recognition, dataset and solver mention extraction, affiliation based geographic mapping and topic modeling to extract fine-grained patterns of data usage that reflect implicit research priorities, data practices, differential reliance on specific databases, evolving modes of data reuse and coupling within scientific workflows, and thematic evolution. Importantly, our proposed methodology is domain-agnostic and transferable to other ORI contexts, and highlights the utility of NLP in quantifying the role of scientific data infrastructures and offers a data-driven reflection on how open-access platforms like LXCat contribute to shaping research directions. This work presents a scalable scientometric framework that has the potential to support evidence based evaluation of ORI platforms and to inform infrastructure design, governance, sustainability, and policy for future development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。