arXiv:2608.04265cs.MAcs.AI2026-08

为智能电网设计了基于物理约束的评估基准,检验大模型规划策略的实际效果。

Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems

  • 构建基于物理过程的可控测试环境,追踪规划指令对系统的影响路径
  • 发现强制搜索策略在所有种子下表现最优,但其他策略存在电压不足恶化问题
  • 提出结合实时可行性判断的优化方法,可将损失降低61.1%且优于固定顺序

当前对大模型规划代理的评估多关注任务是否成功或计划是否被遵循。在战略型网络物理系统中,更关键的问题是:当自主参与者响应且物理规律限制结果时,规划架构是否仍合适。为此,本文构建了一个基于物理约束的受控基准,聚焦于由规划引发的控制轨迹——即执行架构作用于其他代理和物理过程的有序操作与指令序列。该基准在包含40个异构产消者和独立模拟辐射形馈线的智能电网需求响应系统中实现预设的顺序、分层及搜索执行器。大模型仅限于类型化策略声明与短消息指令,而调度构建、产消者动态与潮流计算均保持显式代码。协议采用成对强制模式反事实、共用随机响应抽样及事件级截止时间可行性判定。三个核心发现:(1)架构显著影响结果,强制搜索在全部五个基线种子中均为最优;(2)执行保真度不仅依赖模式一致,客观替代性虽达1.0,却使电压不足增加2.68倍;(3)144场景、576轮实验中,四个架构中有三个生成可行最优解。预设压力测试集的平均遗憾为90.7(95%区间[73.8, 108.6]),无明显优于固定顺序策略;但在质量预测前应用已知截止时间可行性,遗憾降至29.0,并优于固定顺序61.1%。全可行消融实验未超越固定搜索,表明剩余挑战在于可行解内的质量选择。五模型扩展分离出压力条件、状态盲视与不变声明者;延迟尾部分析显示,实时可行性应以概率方式处理。

原文摘要 · Abstract (English)

Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonomous participants respond and physics constrains the outcome. We introduce a controlled, physics-grounded benchmark built around planning-induced control trajectories: the ordered planning operations and directives through which an execution architecture acts on other agents and the physical process. It implements predefined, sequential, hierarchical, and search executors in a smart-grid demand-response system with 40 heterogeneous prosumers and an independently simulated radial feeder. The LLM is bounded to typed policy declaration and short operator messages, while schedule construction, prosumer dynamics, and power flow remain explicit code. The protocol uses paired forced-mode counterfactuals, common random response draws, and event-level deadline feasibility. Three properties follow. Architecture materially changes outcomes: forced search is the oracle in all five baseline seeds. Execution fidelity needs more than mode agreement: objective substitution holds agreement at 1.0 while increasing voltage shortfall by 2.68x. A 144-scenario, 576-episode bank has feasible oracles from three of the four architectures. A prespecified stress-held-out ridge has mean regret 90.7 (95% interval [73.8, 108.6]) and no detectable value over fixed sequential; applying known deadline feasibility before quality prediction cuts regret to 29.0 and improves over fixed sequential by 61.1. An all-feasible ablation does not beat fixed search, localising the remaining challenge to within-feasible quality selection. A five-model extension separates stress-conditioned, state-blind, and invariant declarers; latency tails show that live feasibility should be treated probabilistically.

大模型规划智能电网物理仿真评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。