让菜农实时干预智能温室控温,提升收益并增强算法鲁棒性。
Grower-in-the-Loop Interactive Reinforcement Learning for Greenhouse Climate Control
- 融合菜农经验与强化学习,设计三种人机协同控制算法。
- 策略调整类方法在输入不完美时仍可提升8.4%利润。
- 提出神经网络增强机制,应对菜农输入不完整问题。
温室气候控制直接影响作物生长与资源利用效率。尽管强化学习(RL)在该领域受到关注,但仍面临训练效率低、对初始条件依赖强等问题。交互式强化学习通过结合种植者(菜农)输入与智能体学习,有望解决上述挑战。然而,该方法尚未应用于温室控制,且可能受制于输入不完美。本文探索了在输入不完美条件下应用交互式强化学习的可行性与性能:(1) 针对温室控制特性,开发三种代表性交互式算法(奖励塑造、策略塑造、控制共享);(2) 分析输入特征常相互矛盾,导致菜农输入难以完美;(3) 提出基于神经网络的方法以提升交互式智能体在输入有限时的鲁棒性;(4) 在模拟温室环境中全面评估三类算法表现。结果表明,引入不完美菜农输入的交互式强化学习具有提升智能体性能的潜力。影响动作选择的算法(如策略塑造和控制共享)在输入不完美时分别实现8.4%和6.8%的利润提升;而仅修改奖励函数的奖励塑造算法对不完美输入敏感,导致利润下降9.4%。这凸显了选择合适交互机制的重要性。
原文摘要 · Abstract (English)
Climate control is crucial for greenhouse production as it directly affects crop growth and resource use. Reinforcement learning (RL) has received increasing attention in this field, but still faces challenges, including limited training efficiency and high reliance on initial learning conditions. Interactive RL, which combines human (grower) input with the RL agent's learning, offers a potential solution to overcome these challenges. However, interactive RL has not yet been applied to greenhouse climate control and may face challenges related to imperfect inputs. Therefore, this paper aims to explore the possibility and performance of applying interactive RL with imperfect inputs into greenhouse climate control, by: (1) developing three representative interactive RL algorithms tailored for greenhouse climate control (reward shaping, policy shaping and control sharing); (2) analyzing how input characteristics are often contradicting, and how the trade-offs between them make grower's inputs difficult to perfect; (3) proposing a neural network-based approach to enhance the robustness of interactive RL agents under limited input availability; (4) conducting a comprehensive evaluation of the three interactive RL algorithms with imperfect inputs in a simulated greenhouse environment. The demonstration shows that interactive RL incorporating imperfect grower inputs has the potential to improve the performance of the RL agent. RL algorithms that influence action selection, such as policy shaping and control sharing, perform better when dealing with imperfect inputs, achieving 8.4% and 6.8% improvement in profit, respectively. In contrast, reward shaping, an algorithm that manipulates the reward function, is sensitive to imperfect inputs and leads to a 9.4% decrease in profit. This highlights the importance of selecting an appropriate mechanism when incorporating imperfect inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。