arXiv:2601.02061cs.AIcs.LG2026-01中稿 · NeurIPS

用高阶导数惩罚让强化学习控制更平滑,减少设备损耗

Higher-Order Action Regularization in Deep Reinforcement Learning: From Continuous Control to Building Energy Management

  • 引入三阶导数惩罚(抖动最小化)提升动作平滑性
  • 在建筑温控中使设备开关次数减少60%,节能效果显著
  • 适合关注真实场景部署的智能控制研究者

深度强化学习代理常表现出高频、不稳定的控制行为,导致能耗过高和机械磨损,阻碍实际应用。本文系统研究通过高阶导数惩罚实现动作平滑性正则化,从连续控制基准测试的理论分析,延伸至建筑能源管理的实际验证。在四个连续控制环境中全面评估显示,三阶导数惩罚(抖动最小化)在保持竞争力性能的同时,始终提供最优平滑性。研究进一步拓展至暖通空调(HVAC)控制系统,在该场景下平滑策略使设备切换频率降低60%,带来显著运行效益。本工作确立了高阶动作正则化在能源关键应用中连接强化学习优化与运行约束的有效路径。

原文摘要 · Abstract (English)

Deep reinforcement learning agents often exhibit erratic, high-frequency control behaviors that hinder real-world deployment due to excessive energy consumption and mechanical wear. We systematically investigate action smoothness regularization through higher-order derivative penalties, progressing from theoretical understanding in continuous control benchmarks to practical validation in building energy management. Our comprehensive evaluation across four continuous control environments demonstrates that third-order derivative penalties (jerk minimization) consistently achieve superior smoothness while maintaining competitive performance. We extend these findings to HVAC control systems where smooth policies reduce equipment switching by 60%, translating to significant operational benefits. Our work establishes higher-order action regularization as an effective bridge between RL optimization and operational constraints in energy-critical applications.

强化学习平滑控制能源管理动作正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。