自动化处理科学数据,让原始数据秒变AI可用资产。
Automated Data Readiness for Scientific AI

- 五阶段流水线自动完成数据清洗、转换与结构化
- 跨气候、蛋白质组学等多领域实现数据100%达标
- 支持可复现、可追踪,适合科研团队快速部署
大型科学计算设施管理着海量科学数据,但其在用于人工智能训练前需大量人工转换。现有框架未能统一自动化转换、就绪评估、溯源追踪与智能体调用部署。我们提出 REDI,一个开源框架,采用五阶段统一流程(摄入、预处理、转换、结构化、输出),每阶段均具备可复现性与智能体调用能力;配套工具 SetGo 自动化实现 FAIR 合规与目录发布。在气候、蛋白质组学、材料科学和核聚变领域评估中,REDI 成功将所有原始数据转化为 AI 就绪状态,输出经领域专家验证。初步结果显示,在前沿系统上处理气候数据时,近似理想并行扩展至 100 节点。溯源分析揭示文件 I/O 是主要性能瓶颈,格式选择为关键优化点。结果表明,REDI 可作为跨领域的科学 AI 数据准备平台,将数据处理瓶颈转化为可复用的公共资源。
原文摘要 · Abstract (English)
Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing framework fully unifies automated transformation, readiness assessment, provenance tracking, and agent-native deployment. We present REDI, an open-source framework that addresses this gap through a unified five-stage pipeline (ingest, preprocess, transform, structure, and output) with per-stage instrumentation for reproducibility and deployment as an agent-callable skill; companion tool SetGo automates FAIR compliance and catalog publication. Evaluated across climate, proteomics, materials science, and nuclear fusion, REDI transforms all datasets from raw to AI-ready, with outputs validated against domain-expert references, and preliminary results show near-ideal parallel scaling to 100 nodes on Frontier for the climate case. Provenance-instrumented profiling reveals file I/O as the dominant pipeline cost, with format selection a first-order optimization lever. These results establish REDI as a cross-domain platform providing automated data readiness for scientific AI, transforming data preparation bottlenecks into reproducible, reusable community assets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。