用大模型协调复杂系统,兼顾约束与长期影响。
LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning

- 大模型统筹多单元,任务控制器生成合规动作。
- 连续性感知算法提升长期控制效果,优于传统方法。
- 适合需跨系统协同、约束严苛的工程场景。
在系统交互难建模、操作信息异构且低层动作需满足严格约束的复杂工程系统中,协调多个相互作用单元极具挑战。本文提出一种基于大语言模型(LLM)的分层框架:LLM 根据异构运行上下文协调各单元,而特定任务的控制器或优化器生成可执行且满足约束的动作。进一步引入连续性感知的 GRPO 算法,以捕捉协调决策在后续控制周期中的演化影响。该方法不仅评估即时结果,还考察当前策略下系统的后续发展。我们在多匝道交通控制和虚拟电厂(VPP)能源管理任务上验证了该框架,训练使用简化系统模型,评估采用更真实的模拟器。在两项任务中,所提方法均持续优于直接任务特定控制、优化、端到端强化学习、基于规则及基于强化学习的分层协调,以及仅用提示的 LLM 协调器,证明了异构上下文推理、分层执行和连续性感知策略学习的价值。
原文摘要 · Abstract (English)
Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, operational information is heterogeneous, and low-level actions must satisfy strict constraints. We propose an LLM-based hierarchical framework in which the LLM coordinates interacting units based on heterogeneous operational context, while task-specific controllers or optimizers generate executable and constraint-aware actions. We further introduce Continuation-Aware GRPO to capture the consequences of coordination decisions over subsequent control intervals. Rather than judging a decision only by its immediate outcome, the method also evaluates how the system evolves afterward under the current policy. We validate the framework on multi-ramp traffic control and virtual power plant (VPP) energy management, using simplified system models for training and more realistic simulators for evaluation. Across both tasks, the proposed method consistently outperforms direct task-specific control and optimization, end-to-end reinforcement learning, rule-based and RL-based hierarchical coordination, and prompting-only LLM coordinators, demonstrating the value of heterogeneous-context reasoning, hierarchical execution, and continuation-aware policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。