arXiv:2607.12856cs.LG2026-07被引 1

用可验证奖励微调推理模型,实现建筑温控储能的高效调度。

Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage Control

  • 基于动态规划结果生成密集奖励,通过强化学习微调模型输出逐小时空调设定值。
  • 仅用30个提示训练后,碳排放从70.5降至61.2公斤,接近动态规划最优解60.8公斤。
  • 适合关注建筑能源调度、可解释性强化学习与开放权重模型应用的研究者。

建筑物需根据电网状况调整制冷负荷。热能储存(TES)支持这种调节,但需在存储约束下提前数小时规划。传统模型预测控制(MPC)和强化学习难以在多建筑场景中扩展。本研究采用可验证奖励的强化学习微调(RLVR),将精确离线动态规划(DP)动作价值转化为每个候选动作的密集奖励。仅使用30个训练提示,强化学习微调(RFT)使模型作为上层调度器,从文本状态和预测中输出每小时热泵设定值。评估在简单办公建筑TES基准上进行,此时精确DP可行且最优解已知。RFT将开放权重模型的碳排放从70.5降至61.2公斤,接近DP最优解60.8公斤。GPT-5几乎达到DP和MPC性能,无需任务特定训练;而GPT-4o(非推理型大模型)排放高于无储能基线,表明推理能力至关重要。轨迹分析显示,RFT主要稳定可观测的规划模式(候选比较、前瞻思考、可行性检查),而非创造新策略。鲁棒性与泛化测试表明:强化后的规划模式在预测误差和未见的TES条件下仍有效,并可迁移至电池任务,但结构差异限制了性能提升。基于DP的可验证奖励为开放权重推理模型适配建筑储能调度提供了实用路径。这些结果推动更真实建筑控制测试及城市级能源管理的可扩展验证器发展。

原文摘要 · Abstract (English)

Buildings are expected to shift cooling loads in response to grid conditions. Thermal energy storage (TES) enables this shift, but scheduling it well requires planning hours ahead under storage constraints. Model predictive control (MPC) and reinforcement learning are difficult to scale across buildings. This study instead adapts an open-weight reasoning model through reinforcement learning with verifiable rewards (RLVR). We convert exact offline dynamic-programming (DP) action values into dense rewards for every candidate action. Using only 30 training prompts, reinforcement fine-tuning (RFT) trains the model as an upper-level scheduler that outputs hourly heat-pump setpoints from text-based states and forecasts. Evaluation uses a deliberately simple office-building TES benchmark where exact DP is tractable and the optimum is known. RFT reduces the open-weight model's emissions from 70.5 to 61.2 kg-CO2, close to the DP optimum of 60.8 kg-CO2. GPT-5 nearly matches DP and MPC without task-specific training, while GPT-4o, a non-reasoning LLM, produces higher emissions than the no-storage baseline, so inference-time reasoning appears important. Trace analysis shows that RFT mainly stabilizes observable planning patterns (candidate comparison, look-ahead, and feasibility checking) rather than creating a new strategy. Robustness and generalization tests clarify what transfers: the reinforced planning patterns persist under forecast errors and an unseen TES condition and carry over to a battery task, but its different structure limits the gains. DP-based verifiable rewards offer a practical way to adapt open-weight reasoning models to building storage scheduling. These results motivate higher-fidelity tests of whole-building control and scalable verifiers for city-scale energy management.

能源调度强化学习推理模型碳减排

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。