用图神经网络和大模型预测龙鳞网络应用运行时,提升仿真效率。
SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems
- 融合GNN与LLM,从路由器端口数据中捕捉时空特征。
- 相比传统方法,预测精度显著提升,支持大规模实时仿真。
- 适合高性能计算中的网络性能优化与系统设计者使用。
龙鳞网络凭借高基数、低直径结构,是高性能计算领域的主流互连方案。其主要挑战在于共享链路引发的工作负载干扰。目前常采用并行离散事件仿真(PDES)分析此类干扰,但高保真度的PDES计算开销巨大,难以用于大规模或实时场景。混合仿真通过引入数据驱动的代理模型提供了一种可行替代方案,尤其适用于预测受动态网络流量影响的应用运行时。本文提出 hemodel,结合图神经网络(GNN)与大语言模型(LLM),从端口级路由器数据中同时捕捉空间与时间模式。实验表明, hemodel在准确率上优于现有统计与机器学习基线,可实现高效精准的运行时预测,并支持龙鳞网络的高效混合仿真。
原文摘要 · Abstract (English)
The Dragonfly network, with its high-radix and low-diameter structure, is a leading interconnect in high-performance computing. A major challenge is workload interference on shared network links. Parallel discrete event simulation (PDES) is commonly used to analyze workload interference. However, high-fidelity PDES is computationally expensive, making it impractical for large-scale or real-time scenarios. Hybrid simulation that incorporates data-driven surrogate models offers a promising alternative, especially for forecasting application runtime, a task complicated by the dynamic behavior of network traffic. We present \ourmodel, a surrogate model that combines graph neural networks (GNNs) and large language models (LLMs) to capture both spatial and temporal patterns from port level router data. \ourmodel outperforms existing statistical and machine learning baselines, enabling accurate runtime prediction and supporting efficient hybrid simulation of Dragonfly networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。