arXiv:2511.09725physics.plasm-phcs.LG2025-11

dFL统一处理核聚变数据融合中的对齐、标注与溯源,提升分析效率50倍以上。

The Data Fusion Labeler (dFL): Challenges and Solutions to Data Harmonization, Labeling, and Provenance in Fusion Energy

  • 构建可复现的流程框架,实现多源数据自动对齐与标准化。
  • 支持每小时超200次放电的稳定标注,效率提升50倍以上。
  • 适合核聚变研究者、数据工程师及实时控制团队使用。

核聚变研究日益依赖于整合高分辨率诊断、控制系统和多尺度模拟产生的异构、多模态数据。这些数据体量巨大、结构复杂,亟需新工具实现跨模态系统的系统性对齐与知识提取。本文提出数据融合标注器(dFL),作为统一工作流工具,可在大规模下完成不确定性感知的数据对齐、符合模式的数据融合以及带溯源信息的自动与手动标注。通过将对齐、归一化与标注嵌入可复现且考虑操作顺序的框架中,dFL使分析时间缩短超过50倍(例如,单日标注量从数个提升至每小时超过200次),显著提高标签质量与训练数据可靠性,并实现跨设备可比性。以DIII-D装置为例,dFL已成功应用于自动检测边缘局域模(ELM)与约束区分类,展现出其在数据驱动发现、模型验证与未来点火等离子体实时控制中的核心潜力。

原文摘要 · Abstract (English)

Fusion energy research increasingly depends on the ability to integrate heterogeneous, multimodal datasets from high-resolution diagnostics, control systems, and multiscale simulations. The sheer volume and complexity of these datasets demand the development of new tools capable of systematically harmonizing and extracting knowledge across diverse modalities. The Data Fusion Labeler (dFL) is introduced as a unified workflow instrument that performs uncertainty-aware data harmonization, schema-compliant data fusion, and provenance-rich manual and automated labeling at scale. By embedding alignment, normalization, and labeling within a reproducible, operator-order-aware framework, dFL reduces time-to-analysis by greater than 50X (e.g., enabling >200 shots/hour to be consistently labeled rather than a handful per day), enhances label (and subsequently training) quality, and enables cross-device comparability. Case studies from DIII-D demonstrate its application to automated ELM detection and confinement regime classification, illustrating its potential as a core component of data-driven discovery, model validation, and real-time control in future burning plasma devices.

核聚变数据融合自动化标注智能诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。