arXiv:2411.16707cs.CLcs.AI2024-11被引 48

用多智能体反馈机制提升LLM在电力系统仿真中的表现

Enhancing LLMs for Power System Simulations: A Feedback-driven Multi-agent Framework

  • 设计多智能体框架,融合检索增强、推理优化与错误反馈机制
  • 在两个数据集上达93%~97%成功率,远超主流大模型的30%以下
  • 30秒完成一次仿真,成本仅0.014美元,适合科研高效协作

将实验技术与大语言模型(LLMs)结合正重塑科学研发模式,使AI成为多功能研究助手而非单一解题工具。然而,在电力系统领域,仿真作为关键实验技术,仍受限于LLMs的领域知识不足、推理能力弱及参数处理不精确等问题。为此,本文提出一种反馈驱动的多智能体框架,包含增强型检索增强生成(RAG)、改进推理模块以及带误差反馈的动态环境执行模块。在Daline和MATPOWER的69项多样化任务上验证,该框架分别实现93.13%和96.85%的成功率,显著优于ChatGPT 4o、o1-preview及微调后的GPT-4o(复杂任务成功率均低于30%)。此外,该框架支持快速低成本执行,每次仿真耗时约30秒,平均成本为0.014美元token。整体而言,该可拓展框架为构建面向研究人员的智能LLM助手奠定了基础,助力电力系统研究及其他领域。

原文摘要 · Abstract (English)

The integration of experimental technologies with large language models (LLMs) is transforming scientific research. It positions AI as a versatile research assistant rather than a mere problem-solving tool. In the field of power systems, however, managing simulations -- one of the essential experimental technologies -- remains a challenge for LLMs due to their limited domain-specific knowledge, restricted reasoning capabilities, and imprecise handling of simulation parameters. To address these limitations, this paper proposes a feedback-driven, multi-agent framework. It incorporates three proposed modules: an enhanced retrieval-augmented generation (RAG) module, an improved reasoning module, and a dynamic environmental acting module with an error-feedback mechanism. Validated on 69 diverse tasks from Daline and MATPOWER, this framework achieves success rates of 93.13% and 96.85%, respectively. It significantly outperforms ChatGPT 4o, o1-preview, and the fine-tuned GPT-4o, which all achieved a success rate lower than 30% on complex tasks. Additionally, the proposed framework also supports rapid, cost-effective task execution, completing each simulation in approximately 30 seconds at an average cost of 0.014 USD for tokens. Overall, this adaptable framework lays a foundation for developing intelligent LLM-based assistants for human researchers, facilitating power system research and beyond.

电力系统多智能体反馈机制仿真加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。