arXiv:2607.02771cs.AIcs.CE2026-07被引 1

自动化处理科学数据,让原始数据秒变AI可用资产。

Automated Data Readiness for Scientific AI

论文配图:Automated Data Readiness for Scientific AI
图 1 · 摘自论文原文
  • 五阶段流水线自动完成数据清洗、转换与结构化
  • 跨气候、蛋白质组学等多领域实现数据100%达标
  • 支持可复现、可追踪,适合科研团队快速部署

大型科学计算设施管理着海量科学数据,但其在用于人工智能训练前需大量人工转换。现有框架未能统一自动化转换、就绪评估、溯源追踪与智能体调用部署。我们提出 REDI,一个开源框架,采用五阶段统一流程(摄入、预处理、转换、结构化、输出),每阶段均具备可复现性与智能体调用能力;配套工具 SetGo 自动化实现 FAIR 合规与目录发布。在气候、蛋白质组学、材料科学和核聚变领域评估中,REDI 成功将所有原始数据转化为 AI 就绪状态,输出经领域专家验证。初步结果显示,在前沿系统上处理气候数据时,近似理想并行扩展至 100 节点。溯源分析揭示文件 I/O 是主要性能瓶颈,格式选择为关键优化点。结果表明,REDI 可作为跨领域的科学 AI 数据准备平台,将数据处理瓶颈转化为可复用的公共资源。

原文摘要 · Abstract (English)

Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing framework fully unifies automated transformation, readiness assessment, provenance tracking, and agent-native deployment. We present REDI, an open-source framework that addresses this gap through a unified five-stage pipeline (ingest, preprocess, transform, structure, and output) with per-stage instrumentation for reproducibility and deployment as an agent-callable skill; companion tool SetGo automates FAIR compliance and catalog publication. Evaluated across climate, proteomics, materials science, and nuclear fusion, REDI transforms all datasets from raw to AI-ready, with outputs validated against domain-expert references, and preliminary results show near-ideal parallel scaling to 100 nodes on Frontier for the climate case. Provenance-instrumented profiling reveals file I/O as the dominant pipeline cost, with format selection a first-order optimization lever. These results establish REDI as a cross-domain platform providing automated data readiness for scientific AI, transforming data preparation bottlenecks into reproducible, reusable community assets.

数据工程科学计算AI-ready自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。