arXiv:2606.20950cs.AIcs.SY2026-06被引 1

为电力工程智能体设计可执行评估基准,验证其操作真实性和可行性。

Power Systems Agent Benchmark: Executable Evaluation of AI Agents in Electric Power Engineering

  • 用程序自动验证智能体动作后果,而非仅评估文字描述。
  • 覆盖8个电力领域41类任务,每项基于可追溯的工程标准。
  • 支持私有生成测试用例,防止数据污染,适合研究与工业级评估。

可执行评估——通过程序检查智能体行为后果而非评分其文本输出——已成为软件场景中评估工具使用型智能体的重要方式。电力工程领域尚无类似基准:当前语言模型应用仍以检索和问答为主,而能操作电力系统实体的智能体多停留在学术原型阶段。本文提出电力系统智能体基准(Power Systems Agent Benchmark),一个面向电力工程智能体的可执行评估框架。智能体接收结构化任务并返回结构化解法;确定性评估器重新计算工程量、校验运行约束,并返回可行性标志、归一化得分及具体违规项。该基准涵盖电力系统八个领域的41类任务,包括潮流分析、继电保护、稳定性、微电网、可靠性、电能质量与预测等,每项任务均基于可引用的文献、标准或明确的工程公式。为防止数据污染,保留的测试用例由各任务族私有种子动态生成,构造过程可审计但实例保持私密。参考评估中,三个命令行智能体表现一致:最强模型接近紧凑层级上限,小型开源模型稍逊,公开与保留集表现基本吻合;另一组独立公开分割网格测试进一步揭示了工具调用效应。该评估同时充当质量控制机制:一致失败可暴露候选任务或评估器缺陷,曾发现被自洽性检查遗漏的潜在评估器漏洞。评估器为紧凑且确定性的代理,但任务契约允许未来升级为基于仿真器的验证,无需改变任务表达或求解方式。

原文摘要 · Abstract (English)

Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way to assess tool-using AI agents in software settings. Electric power engineering has not yet had an analogous benchmark: language-model use is still dominated by retrieval and text question answering, while agents acting on power-system artifacts remain mostly academic prototypes. We introduce the Power Systems Agent Benchmark, an executable benchmark for power-engineering agents. An agent receives a structured task and returns a structured solution; a deterministic evaluator recomputes the engineering quantities, checks operational constraints, and returns a feasibility flag, a normalized score, and explicit violations. The benchmark contains 41 task families across eight areas of power engineering, from power flow and protection to stability, microgrids, reliability, power quality, and forecasting. Each task is grounded in a citable source, standard, or documented engineering formulation. To resist contamination, held-out cases are synthesized on demand by per-family generators from private seeds: the construction is inspectable, but the instances remain private. In a reference evaluation with three command-line agents, the strongest score near the compact tier's ceiling, a smaller open model trails, and public and held-out performance are broadly consistent; a separate public-split grid with OpenCode and Aider probes harness effects. The reference evaluation doubles as quality control: unanimous failures flag candidate task or evaluator defects, and it exposed a latent evaluator bug missed by self-consistency checks. The evaluators are compact deterministic surrogates, but the task contract allows their internals to be upgraded to simulator-backed checks without changing how tasks are posed or solved.

电力系统智能体评估可执行验证基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。