用大模型解决科研软件元数据的身份识别问题,提升数据整合效率。
Identity resolution of software metadata using Large Language Models
- 采用指令微调的大语言模型进行软件元数据身份匹配
- 代理方法实现高精度自动化判断,准确率显著优于基准
- 适用于科研软件可持续性分析与开放平台的数据治理
软件是科研的重要组成部分,但其元数据管理长期被忽视。来自bio.tools、Bioconductor和Galaxy ToolShed等平台的结构化元数据为生命科学领域的研究软件提供了宝贵信息。尽管这些元数据最初用于发现与集成,现可拓展用于大规模软件实践分析。然而,不同平台间元数据质量与完整性差异显著,反映多样化的文档规范。为全面理解软件开发与可持续性,需整合这些元数据,但面临异构性和规模挑战。本文评估了多种指令微调的大语言模型在软件元数据身份识别任务中的表现,该任务是构建统一科研软件集合的关键步骤。该集合是OpenEBench上Software Observatory的核心参考数据。我们基于人工标注的标准进行模型对比,分析模糊案例表现,并提出一种基于共识的高置信度自动化决策代理。该代理在精度与统计稳健性上表现优异,同时揭示当前模型局限及跨注册库与存储库实现FAIR对齐元数据语义判断的普遍挑战。
原文摘要 · Abstract (English)
Software is an essential component of research. However, little attention has been paid to it compared with that paid to research data. Recently, there has been an increase in efforts to acknowledge and highlight the importance of software in research activities. Structured metadata from platforms like bio.tools, Bioconductor, and Galaxy ToolShed offers valuable insights into research software in the Life Sciences. Although originally intended to support discovery and integration, this metadata can be repurposed for large-scale analysis of software practices. However, its quality and completeness vary across platforms, reflecting diverse documentation practices. To gain a comprehensive view of software development and sustainability, consolidating this metadata is necessary, but requires robust mechanisms to address its heterogeneity and scale. This article presents an evaluation of instruction-tuned large language models for the task of software metadata identity resolution, a critical step in assembling a cohesive collection of research software. Such a collection is the reference component for the Software Observatory at OpenEBench, a platform that aggregates metadata to monitor the FAIRness of research software in the Life Sciences. We benchmarked multiple models against a human-annotated gold standard, examined their behavior on ambiguous cases, and introduced an agreement-based proxy for high-confidence automated decisions. The proxy achieved high precision and statistical robustness, while also highlighting the limitations of current models and the broader challenges of automating semantic judgment in FAIR-aligned software metadata across registries and repositories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。