arXiv:2601.08052cs.AI2026-01

用预测感知强化学习优化奶牛场用电,降本增效更稳定

Forecast Aware Deep Reinforcement Learning for Efficient Electricity Load Scheduling in Dairy Farms

  • 引入小时与月份校准的短期预测,动态调整用电策略
  • 相比PPO降电费1%,比DQN降4.8%,电池调峰减少13.1%电网依赖
  • 适合关注能源调度、可再生能源整合的农业与能源研究者

奶牛养殖是高耗能行业,严重依赖电网电力。随着可再生能源接入增加,可持续能源管理对降低电网依赖、实现联合国可持续发展目标7(可负担的清洁能源)至关重要。然而,可再生能源的间歇性给供需实时平衡带来挑战。智能负荷调度因此成为降低运营成本并保障可靠性的重要手段。强化学习在提升能效、降低成本方面展现潜力,但多数基于强化学习的调度方法假设未来电价或发电量完全可知,这在动态环境中不现实。此外,标准PPO变体依赖固定剪裁或KL散度阈值,常导致变量电价下训练不稳定。为此,本文提出一种面向奶牛场高效用电调度的深度强化学习框架,聚焦电池储能与热水加热,在真实运营约束下实现优化。所提预测感知PPO通过基于小时和月份的残差校准,融合短期需求与可再生能源生成预测;而PID KL PPO采用比例积分微分控制器动态调节KL散度,实现稳定策略更新。基于真实奶牛场数据训练,该方法相比标准PPO降低电费1%,比DQN低4.8%,比SAC低1.5%。在电池调度中,该方法使电网购电减少13.1%,展现出在现代奶牛场可持续能源管理中的可扩展性与有效性。

原文摘要 · Abstract (English)

Dairy farming is an energy intensive sector that relies heavily on grid electricity. With increasing renewable energy integration, sustainable energy management has become essential for reducing grid dependence and supporting the United Nations Sustainable Development Goal 7 on affordable and clean energy. However, the intermittent nature of renewables poses challenges in balancing supply and demand in real time. Intelligent load scheduling is therefore crucial to minimize operational costs while maintaining reliability. Reinforcement Learning has shown promise in improving energy efficiency and reducing costs. However, most RL-based scheduling methods assume complete knowledge of future prices or generation, which is unrealistic in dynamic environments. Moreover, standard PPO variants rely on fixed clipping or KL divergence thresholds, often leading to unstable training under variable tariffs. To address these challenges, this study proposes a Deep Reinforcement Learning framework for efficient load scheduling in dairy farms, focusing on battery storage and water heating under realistic operational constraints. The proposed Forecast Aware PPO incorporates short term forecasts of demand and renewable generation using hour of day and month based residual calibration, while the PID KL PPO variant employs a proportional integral derivative controller to regulate KL divergence for stable policy updates adaptively. Trained on real world dairy farm data, the method achieves up to 1% lower electricity cost than PPO, 4.8% than DQN, and 1.5% than SAC. For battery scheduling, PPO reduces grid imports by 13.1%, demonstrating scalability and effectiveness for sustainable energy management in modern dairy farming.

能源调度强化学习可再生能源农业能源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。