arXiv:2504.05307cs.IRcs.AI2025-04被引 5

用AI自动修复科研元数据,让数据更易被找到。

Toward Total Recall: Enhancing FAIRness through AI-Driven Metadata Standardization

  • 用GPT-4+CEDAR模板批量修正元数据格式与内容
  • 元数据召回率从17.65%提升至62.87%
  • 适合数据管理、生物信息学研究者使用

科学元数据常存在不完整、不一致和格式错误,阻碍数据发现与重用。本文提出一种结合GPT-4与CEDAR知识库结构化模板的方法,自动标准化元数据并确保符合标准。CEDAR模板定义了元数据提交的必填字段及其允许值。通过模板引导GPT-4批量修正元数据,显著提升检索性能,尤其在召回率上表现突出——即从全部相关数据中成功检索出的比例。基于美国国家生物技术信息中心(NCBI)维护的BioSample和GEO数据库的实验表明,经GPT-4+CEDAR处理的元数据检索效果远优于原始状态及仅依赖数据词典指导的GPT-4+DD方案。平均召回率从基准原始元数据的17.65%大幅提升至62.87%。此外,对比LLaMA-3和MedLLaMA2等其他大模型,GPT-4+CEDAR仍保持显著优势。结果表明,将先进语言模型与标准化元数据结构相结合,可显著提升数据检索效率与可靠性,加速科学发现与数据驱动研究。

原文摘要 · Abstract (English)

Scientific metadata often suffer from incompleteness, inconsistency, and formatting errors, which hinder effective discovery and reuse of the associated datasets. We present a method that combines GPT-4 with structured metadata templates from the CEDAR knowledge base to automatically standardize metadata and to ensure compliance with established standards. A CEDAR template specifies the expected fields of a metadata submission and their permissible values. Our standardization process involves using CEDAR templates to guide GPT-4 in accurately correcting and refining metadata entries in bulk, resulting in significant improvements in metadata retrieval performance, especially in recall -- the proportion of relevant datasets retrieved from the total relevant datasets available. Using the BioSample and GEO repositories maintained by the National Center for Biotechnology Information (NCBI), we demonstrate that retrieval of datasets whose metadata are altered by GPT-4 when provided with CEDAR templates (GPT-4+CEDAR) is substantially better than retrieval of datasets whose metadata are in their original state and that of datasets whose metadata are altered using GPT-4 with only data-dictionary guidance (GPT-4+DD). The average recall increases dramatically, from 17.65\% with baseline raw metadata to 62.87\% with GPT-4+CEDAR. Furthermore, we evaluate the robustness of our approach by comparing GPT-4 against other large language models, including LLaMA-3 and MedLLaMA2, demonstrating consistent performance advantages for GPT-4+CEDAR. These results underscore the transformative potential of combining advanced language models with symbolic models of standardized metadata structures for more effective and reliable data retrieval, thus accelerating scientific discoveries and data-driven research.

元数据AI修复数据发现生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。