arXiv:2607.11915physics.plasm-phcs.LG2026-07

首次系统评估等离子体诊断模型在传感器故障下的鲁棒性,发现临界时刻故障会致命影响序列模型。

Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark

论文配图:Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark
图 1 · 摘自论文原文
  • 基于TokaMark数据集,在6种故障场景下测试4类模型的鲁棒性
  • 临界时刻故障使LSTM性能恶化212%而XGBoost仅增37%
  • 仅均值填充可恢复报警准确率至TPR=1.00,适合关键预警场景

托卡马克聚变装置的等离子体诊断模型通常在干净完整的传感器数据上评估。实际上,诊断系统常出现故障:采集系统延迟启动、单个传感器失效,且信号丢失恰好发生在等离子体破裂前。我们首次使用包含11,573次MAST放电的TokaMark数据集,对XGBoost、LSTM、Transformer及TokaMark CNN基线模型,在六种物理驱动的故障场景和三种插补策略下进行系统性鲁棒性评估。提出鲁棒性评分(RS)实现跨架构标准化比较。核心发现:破裂临近时的传感器故障(最后时间窗口注入干扰)使序列模型性能严重下降(LSTM NRMSE增加212%),而统计特征模型相对稳定(XGBoost仅增加37%)。前向填充可消除随机丢包对序列模型的大部分影响(从+57%降至接近0%),但在末尾窗口故障时帮助甚微。基于真实破裂时间戳的射线级报警评估显示,临界故障下LSTM报警检测召回率暴跌至TPR=0.00,而均值填充使其恢复至TPR=1.00,与NRMSE结果呈现相反模式。等离子体电流在所有架构中均为最关键诊断指标(移除后性能下降73%~140%)。代码、数据与训练检查点已开源。

原文摘要 · Abstract (English)

Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. In practice, fusion diagnostics fail regularly: acquisition systems start late, individual sensors die, and signal dropouts cluster precisely when a plasma disruption is approaching. We present the first systematic robustness benchmark for plasma diagnostic ML using the TokaMark dataset of 11,573 MAST shots, evaluating XGBoost, LSTM, Transformer, and the TokaMark CNN baseline across six physically-grounded failure scenarios and three imputation strategies. We introduce the Robustness Score (RS) for standardized cross-architecture comparison. Our central finding is that disruption-proximate sensor failure (corruption injected in the final window timesteps) collapses sequence model performance (LSTM +212% NRMSE) while a statistical feature model remains comparatively stable (XGBoost +37%). Forward-fill imputation eliminates nearly all degradation from random dropout for sequence models (LSTM +57% to ~0%), but offers little help when the end of the window is corrupted. Shot-level alarm evaluation using ground-truth disruption timestamps reveals that LSTM alarm detection collapses to TPR=0.00 under proximate sensor failure, while mean-fill imputation recovers it to TPR=1.00, a reversal of the pattern observed in NRMSE. Plasma current emerges as the single most critical diagnostic across all architectures (+73% to +140% upon removal). Code, data, and trained checkpoints are available at https://github.com/Neerav-Gupta/tokamark-robustness.

等离子体诊断模型鲁棒性传感器故障机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。