用实时查询知识库的AI自动规范生物医学元数据,提升数据可重用性。
Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent
- AI Agent动态调用标准指南和术语库,实时获取准确规范
- 在839条旧元数据上,准确率显著优于纯LLM方法
- 适合需要提升数据合规性的科研团队和数据平台
科学元数据常不完整且不符合社区标准,限制数据的可发现性、互操作性和重用性。尽管已有标准报告指南,但通常缺乏机器可操作性。构建符合FAIR原则的数据集需将标准转化为带丰富字段说明和精确取值约束的机器可读模板。近期研究显示,基于字段名和本体约束引导的LLM能改善元数据标准化,但这些方法将约束视为静态提示,仅依赖模型自身训练知识。本文提出一种基于LLM的元数据标准化系统,通过实时查询标准报告指南和权威生物医学术语服务,按需获取规范标准。我们在人类生物分子图谱计划(HuBMAP)的839条遗留元数据上进行评估,采用专家标注的黄金标准进行精确匹配评估。结果表明,在包含和不包含本体约束的字段上,引入实时工具访问的LLM在预测准确性上均持续优于仅使用LLM的方法,证明了该方法在自动化生物医学元数据标准化中的实用性。
原文摘要 · Abstract (English)
Scientific metadata are often incomplete and noncompliant with community standards, limiting dataset findability, interoperability, and reuse. Even when standard metadata reporting guidelines exist, they typically lack machine-actionable representations. Producing FAIR datasets requires encoding metadata standards as machine-actionable templates with rich field specifications and precise value constraints. Recent work has shown that LLMs guided by field names and ontology constraints can improve metadata standardization, but these approaches treat constraints as static text prompts, relying on the model's training knowledge alone. We present an LLM-based metadata standardization system that queries standard reporting guidelines and authoritative biomedical terminology services in real time to retrieve canonically correct standards on demand. We evaluate this approach on 839 legacy metadata records from the Human BioMolecular Atlas Program (HuBMAP) using an expert-curated gold standard for exact-match assessment. Our evaluation shows that augmenting the LLM with real-time tool access consistently improves prediction accuracy over the LLM alone across both ontology-constrained and non-ontology-constrained fields, demonstrating a practical approach to automated standardization of biomedical metadata.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。