通过集成方法实现托卡马克约束态分类的不确定性量化与鲁棒性提升
Robust Confinement State Classification with Uncertainty Quantification through Ensembled Data-Driven Methods
- 采用动态与静态模型集成,结合多类诊断特征进行分类
- 在302次TCV放电中实现高精度识别,L/D/H模态分类κ系数达0.89
- 可处理信号缺失或损坏场景,输出可信度估计,适合实时控制应用
托卡马克聚变性能依赖于高能量约束,通常通过不同运行模式实现。自动标注这些约束态对大规模分析和实时控制至关重要。尽管在状态转换或边缘情形下自动化难度增加,数据驱动模型已取得显著进展。但现有方法多提供点估计,难以应对缺失或损坏的输入信号。为此,我们提出一种具有不确定性量化和模型鲁棒性的约束态分类方法。聚焦于TCV放电的离线分析,区分L模、H模及中间振荡相(D)。通过双轴集成:模型架构(基于循环傅里叶神经算子的动态模型与基于梯度提升决策树的静态模型)和特征集(按诊断系统或物理量划分)。构建了302次完整标注的TCV放电数据集,将公开发布。使用Cohen's kappa系数评估预测性能,用期望校准误差评估不确定性校准效果。同时分析常见与特殊场景、各组件表现、分布外泛化能力、信号缺失/损坏情况,以及不同状态转换附近的条件平均行为。结果表明,该方法能高精度区分L、D、H模态,具备强鲁棒性,并提供有意义的不确定性估计。
原文摘要 · Abstract (English)
Maximizing fusion performance in tokamaks relies on high energy confinement, often achieved through distinct operating regimes. The automated labeling of these confinement states is crucial to enable large-scale analyses or for real-time control applications. While this task becomes difficult to automate near state transitions or in marginal scenarios, much success has been achieved with data-driven models. However, these methods generally provide predictions as point estimates, and cannot adequately deal with missing and/or broken input signals. To enable wide-range applicability, we develop methods for confinement state classification with uncertainty quantification and model robustness. We focus on off-line analysis for TCV discharges, distinguishing L-mode, H-mode, and an in-between dithering phase (D). We propose ensembling data-driven methods on two axes: model formulations and feature sets. The former considers a dynamic formulation based on a recurrent Fourier Neural Operator-architecture and a static formulation based on gradient-boosted decision trees. These models are trained using multiple feature groupings categorized by diagnostic system or physical quantity. A dataset of 302 TCV discharges is fully labeled, and will be publicly released. We evaluate our method quantitatively using Cohen's kappa coefficient for predictive performance and the Expected Calibration Error for the uncertainty calibration. Furthermore, we discuss performance using a variety of common and alternative scenarios, the performance of individual components, out-of-distribution performance, cases of broken or missing signals, and evaluate conditionally-averaged behavior around different state transitions. Overall, the proposed method can distinguish L, D and H-mode with high performance, can cope with missing or broken signals, and provides meaningful uncertainty estimates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。