arXiv:2602.10132physics.plasm-phcs.AI2026-02被引 2

构建融合等离子体模型评估基准,推动人工智能在核聚变中的应用。

TokaMark: A Comprehensive Benchmark for MAST Tokamak Plasma Models

  • 基于MAST实验数据构建统一评估框架
  • 涵盖14项任务覆盖多种物理机制与诊断手段
  • 开源数据与工具,适合核聚变与AI交叉研究者

商业化聚变反应堆如托卡马克的开发与运行依赖于从稀疏、噪声大且不完整的传感器读数中准确预测等离子体动态。传统数值方法难以应对底层物理复杂性及实验数据异质性,凸显现代数据原生方法的潜力。然而,其发展受限于缺乏经过整理的公开数据集和标准化评估基准。现有聚变数据集稀缺、分散于不同机构、具有设施特异性且标注不一致,影响可复现性,阻碍了人工智能方法的公平、可扩展比较。本文提出TokaMark,一个针对大型球形托卡马克(MAST)真实实验数据的综合性评估基准。TokaMark提供一套统一的多模态聚变数据访问与标准化评估协议工具集,包含14项涵盖多种物理机制的任务,利用多样化诊断手段并覆盖多个应用场景。同时提供基线模型,支持在统一框架内进行透明比较与验证。通过建立统一基准,TokaMark旨在加速数据驱动的人工智能等离子体建模进展,助力实现可持续稳定的聚变能源。数据集、基准、文档与工具已开源,地址:https://github.com/UKAEA-IBM-STFC-Fusion-FMs/tokamark_baseline。

原文摘要 · Abstract (English)

Development and operation of commercially viable fusion energy reactors such as tokamaks require accurate predictions of plasma dynamics from sparse, noisy, and incomplete sensors readings. The complexity of the underlying physics and the heterogeneity of experimental data pose formidable challenges for conventional numerical methods, and highlight the promise of modern data-native approaches. A major obstacle in realizing this potential is, however, the lack of curated, openly available datasets and standardized benchmarks. Existing fusion datasets are scarce, fragmented across institutions, facility-specific, and inconsistently annotated, which limits reproducibility and prevents a fair and scalable comparison of AI approaches. In this paper, we introduce TokaMark, a structured benchmark to evaluate AI models on real experimental data collected from the Mega Ampere Spherical Tokamak (MAST). TokaMark provides a comprehensive suite of tools designed to unify access to multi-modal fusion data and standardize evaluation protocols. The benchmark includes a curated list of 14 tasks spanning a range of physical mechanisms, exploiting a variety of diagnostics and covering multiple operational use cases. A baseline model is provided to facilitate transparent comparison and validation within a unified framework. By establishing a unified benchmark, TokaMark aims to accelerate progress in data-driven AI-based plasma modeling, contributing to the broader goal of achieving sustainable and stable fusion energy. The dataset, benchmark, documentation, and tooling are open-sourced under https://github.com/UKAEA-IBM-STFC-Fusion-FMs/tokamark_baseline.

核聚变等离子体建模AI基准数据驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。