为科学大模型训练构建数据就绪评估框架,助力跨领域高效AI研发。
Data Readiness for Scientific AI at Scale
- 提出二维数据就绪框架,涵盖从原始到AI可用的数据处理阶段。
- 分析气候、核聚变等四大领域,提炼通用预处理模式与特殊约束。
- 聚焦高性能计算环境,支持可复现的大规模科学AI训练。
本文研究数据就绪于人工智能(DRAI)原则在领导级科学数据集中的应用,这些数据集用于训练基础模型。我们分析了气候、核聚变、生物/健康和材料四个代表性领域的典型工作流,识别出共性的预处理模式与领域特定约束。提出一个二维就绪框架,包含数据就绪等级(从原始到AI就绪)与数据处理阶段(从摄入到分片),均适配高性能计算(HPC)环境。该框架揭示了将科学数据转化为可扩展AI训练资源的关键挑战,尤其关注基于Transformer的生成模型。两个维度共同构成概念成熟度矩阵,用于刻画科学数据就绪状态,并指导基础设施建设,实现跨领域标准化、可复现的大规模科学AI支持。
原文摘要 · Abstract (English)
This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains - climate, nuclear fusion, bio/health, and materials - to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework composed of Data Readiness Levels (raw to AI-ready) and Data Processing Stages (ingest to shard), both tailored to high performance computing (HPC) environments. This framework outlines key challenges in transforming scientific data for scalable AI training, emphasizing transformer-based generative models. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。