用AI自动识别论文里的研究软件,助其成为可发现、可引用的学术资产
Making Software FAIR: A machine-assisted workflow for the research software lifecycle
- 基于机器学习从论文中自动识别研究软件资源
- 通过作者验证后为软件注册持久标识符并归档
- 适合科研人员和机构提升软件可重用性
当前开放研究软件常隐藏在论文正文内,难以被发现、引用和复用。为使这些资源成为首等文献记录,需先识别并注册持久标识符(PIDs),以符合FAIR原则(可发现、可访问、可互操作、可重用)。目前多数研究软件仍不满足该标准,且极少与发表论文明确关联。SoFAIR是2024-2025年为期两年的国际项目,依托全球开放存储库网络,扩展CORE、Software Heritage、HAL等基础设施及GROBID工具的能力,实现研究软件全生命周期管理:1)利用机器学习从学术论文中自动识别研究软件资产;2)由作者验证识别结果;3)为软件资产注册PID并进行存档。
原文摘要 · Abstract (English)
A key issue hindering discoverability, attribution and reusability of open research software is that its existence often remains hidden within the manuscript of research papers. For these resources to become first-class bibliographic records, they first need to be identified and subsequently registered with persistent identifiers (PIDs) to be made FAIR (Findable, Accessible, Interoperable and Reusable). To this day, much open research software fails to meet FAIR principles and software resources are mostly not explicitly linked from the manuscripts that introduced them or used them. SoFAIR is a 2-year international project (2024-2025) which proposes a solution to the above problem realised over the content available through the global network of open repositories. SoFAIR will extend the capabilities of widely used open scholarly infrastructures (CORE, Software Heritage, HAL) and tools (GROBID) operated by the consortium partners, delivering and deploying an effective solution for the management of the research software lifecycle, including: 1) ML-assisted identification of research software assets from within the manuscripts of scholarly papers, 2) validation of the identified assets by authors, 3) registration of software assets with PIDs and their archival.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。