构建核聚变多模态数据组织框架,解决科学领域数据异构难题。
Towards Large-Scale Heterogeneous Data Organization for Scientific Foundation Models: A Nuclear Fusion Case Study

- 整合20余种传感器、跨5个数量级采样率的异构数据
- 支持点测量、光谱图、图像等混合张量结构建模
- 为多模态控制与核聚变研究提供可扩展数据范式
训练高效基础模型需要大规模且有序的数据集,但核聚变等科学领域因数据高度异构且稀疏而面临独特挑战。本文分析了用于构建此类模型的数据:包含超过20种传感器类型,采样率跨度达5个数量级,数据结构混合(点测量、光谱图、图像),且物理过程具有非平稳性。我们评估了输入复杂性,探讨了时间上下文与频率分辨率之间的权衡。该分析为大规模多模态波动数据的表示提供了模板,对多模态控制系统及核聚变研究均具重要意义。
原文摘要 · Abstract (English)
Training effective foundation models requires massive and organized datasets, yet scientific domains such as nuclear fusion present unique challenges due to largely heterogeneous and sparse data. Here we characterize the data used in developing such a model: with over 20 sensor types spanning 5 orders of magnitude in sampling rate, mixed tensor structures (point measurements, spectrograms, images), and nonstationary physics. We analyze our input complexity and discuss trade-offs between temporal context and frequency resolution. Our analysis provides a template for representing multi-modal fluctuation data at scale, with implications for both multi-modal control systems and nuclear fusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。