SetGo自动检测并修复科学数据元数据缺陷,提升可发现性和可用性。
SetGo: Metadata Readiness for Scientific AI Datasets
- 基于六维度评估元数据完整性、合规性与治理状态,支持修复与发布。
- 修复后数据的FAIR评分从52-57%提升至81-91%,显著改善可重用性。
- 兼容LLM智能代理,支持自然语言操作,降低科研人员使用门槛。
面向AI的科学数据集不仅需要适合模型训练的计算就绪性,还需具备利于发现、共享与复用的元数据就绪性。现有工具REDI解决计算就绪性问题,但缺乏对元数据完整、合规与标准符合性的评估。现有FAIR评估仅针对已发布的仓库记录,且无统一系统覆盖FAIR合规、许可、溯源、治理、可复现性与目录就绪性。本文提出SetGo,一个开源Python工具包,可在数据发布或归档前评估并修复这六个维度的元数据就绪性。应用于四个科学语料库显示:ERA5气候数据在ACDD 1.3标准下仅得4%合规分;材料数据不符合OPTIMADE物种定义要求;由PDB衍生的蛋白质组学数据存在与SPDX标识符不兼容的许可条款。经引导式增强,整体FAIR得分从52-57%提升至81-91%。单个setgo publish命令可将数据推送至Hugging Face Hub、CKAN或OpenMetadata,并附带ML Commons Croissant 1.0元数据侧车文件。通过/ setgo技能集成大语言模型驱动的编码代理,支持交互与自动化工作流,用户仅需提供缺失的元数据值即可完成全链路评估-增强-发布。
原文摘要 · Abstract (English)
Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset's metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52-57% to 81-91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess-enrich-publish loop, with user involvement limited to supplying missing metadata values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。