测试大模型在电力优化约束下的推理与求解能力
Can Large Language Models Reason and Optimize Under Constraints?
- 构建电力系统最优潮流约束任务,评估模型多技能协同能力
- 主流大模型在多数任务中表现失败,复杂场景下推理模型仍不奏效
- 为电网优化类应用提供严格评测基准,适合关注可信AI的开发者
大型语言模型(LLMs)在自然语言任务中展现出强大能力,但其在存在约束的抽象与优化问题上的表现尚未被充分探索。本文研究大模型是否能在最优潮流(OPF)问题的物理和运行约束下进行推理与优化。我们设计了一项具有挑战性的评估任务,要求模型具备推理、结构化输入处理、算术运算及约束优化等核心能力。评估结果表明,当前最先进的大模型在大多数任务中表现失败,即使是专门强化推理能力的模型,在最复杂场景下依然无法完成任务。研究揭示了大模型在结构化约束推理方面存在显著能力缺口,并为开发能解决真实电网优化问题的高可靠大模型助手提供了严格的测试环境。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated great capabilities across diverse natural language tasks; yet their ability to solve abstraction and optimization problems with constraints remains scarcely explored. In this paper, we investigate whether LLMs can reason and optimize under the physical and operational constraints of Optimal Power Flow (OPF) problem. We introduce a challenging evaluation setup that requires a set of fundamental skills such as reasoning, structured input handling, arithmetic, and constrained optimization. Our evaluation reveals that SoTA LLMs fail in most of the tasks, and that reasoning LLMs still fail in the most complex settings. Our findings highlight critical gaps in LLMs' ability to handle structured reasoning under constraints, and this work provides a rigorous testing environment for developing more capable LLM assistants that can tackle real-world power grid optimization problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。