EPIC用多智能体架构加速高性能计算数据分析,降本提效。
EPIC: Generative AI Platform for Accelerating HPC Operational Data Analytics
- 分层多智能体设计,大模型统筹小模型分工处理
- 小模型在描述性分析上比大模型高26%准确率
- 混合方案降低19倍大模型使用成本,适合运维人员
我们提出EPIC,一个由人工智能驱动的平台,用于增强高性能计算(HPC)的运行数据解析。EPIC采用分层多智能体架构,顶层大型语言模型负责查询处理、推理与综合,协调三个低层级专用智能体完成信息检索、描述性分析和预测性分析。该架构使EPIC能够动态、迭代地处理文本、图像和表格等多种模态数据。针对现有静态分析方法难以适应任务演化与用户需求的问题,我们在前沿(Frontier)HPC系统上进行了全面评估。结果表明,以描述性分析为例,经过微调的小模型性能优于当前主流的大规模基础模型,准确率最高提升26%;同时,通过结合大模型与微调后的本地开源权重模型的混合策略,相比专有解决方案,实现了19倍的大型语言模型运营成本节约。
原文摘要 · Abstract (English)
We present EPIC, an AI-driven platform designed to augment operational data analytics. EPIC employs a hierarchical multi-agent architecture where a top-level large language model provides query processing, reasoning and synthesis capabilities. These capabilities orchestrate three specialized low-level agents for information retrieval, descriptive analytics, and predictive analytics. This architecture enables EPIC to perform HPC operational analytics on multi-modal data, including text, images, and tabular formats, dynamically and iteratively. EPIC addresses the limitations of existing HPC operational analytics approaches, which rely on static methods that struggle to adapt to evolving analytics tasks and stakeholder demands. Through extensive evaluations on the Frontier HPC system, we demonstrate that EPIC effectively handles complex queries. Using descriptive analytics as a use case, fine-tuned smaller models outperform large state-of-the-art foundation models, achieving up to 26% higher accuracy. Additionally, we achieved 19x savings in LLM operational costs compared to proprietary solutions by employing a hybrid approach that combines large foundational models with fine-tuned local open-weight models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。