用语义引导的扩散模型构建高性能计算的动态数字孪生
SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series
- 以传感器语义描述为条件生成时间序列,解耦系统与数据结构
- 重建任务MAE达0.0470,热推断任务MAE低至0.033
- 零样本排列稳定,单模型可完成补全、预测、虚拟传感
数据中心及其计算节点需要能够精确建模工作负载、环境参数与物理指标复杂交互的动态数字孪生。当前用于高性能计算(HPC)及遥测数据的机器学习方法通常依赖于固定位置、匿名的静态传感器变量集,仅适用于单一任务,一旦任务或传感器指标变化即失效。我们提出SeT-Diff,首个面向计算节点遥测与时间序列的基础模型。不同于僵化架构,其基于扩散模型,以每个传感器的语义描述作为生成条件,将系统动态与数据结构解耦。在真实超算数据集上的实验表明,重建任务的平均绝对误差(MAE)为0.0470。SeT-Diff具备零样本排列稳定性,在传感器顺序随机打乱时仍保持精度几乎无损。单一预训练模型可有效完成数据补全、预测与虚拟传感,热推断任务中达到0.033的MAE,成为高效的数据驱动式HPC数字孪生。
原文摘要 · Abstract (English)
Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, environmental parameters, and physical metrics. Current machine learning approaches for HPC and its telemetry typically rely on a static subset of anonymous, fixed-position sensor variables tailored to single tasks. Consequently, these models become obsolete when target tasks change or sensor metrics vary. We propose SeT-Diff, the first foundational model for compute node telemetry and time-series. Unlike rigid architectures, our diffusion-based approach conditions the generative process on each sensor's semantic description, decoupling the system dynamics from the structure of the dataset. Experiments on a real-world supercomputer dataset demonstrate a Mean Absolute Error (MAE) of 0.0470 on reconstruction tasks. SeT-Diff exhibits zero-shot permutation stability, maintaining accuracy with negligible degradation even when sensors are shuffled. A single pre-trained model effectively performs data imputation, forecasting, and virtual sensing - achieving a 0.033 MAE in thermal inference - making SeT-Diff an effective data-driven digital twin for HPC systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。