arXiv:2607.11959cs.AI2026-07

为智能温室强化学习控制提供可复现的奖励组件审计框架

Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses

论文配图:Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses
图 1 · 摘自论文原文
  • 先校准后审计,分解奖励为温、湿、光等具体控制成分
  • 在多个场景下保持奖励组件可比性,包括模拟训练与真实数据
  • 适合温室控制工程师和算法研究者评估策略实际表现

智能温室强化学习可快速大规模测试气候调控方案,但仅依赖单一模拟回报不足:还需明确策略何时加热、增补CO2、通风、调控湿度、部署遮阳帘或使用补光灯。本文提出一种可复现的校准优先奖励组件审计框架,使温控、CO2、湿度、蒸气压差、遮阳帘及执行动作代理等命名组件在模拟训练、设施适配回放、已记录的自主温室挑战数据以及执行器规则提炼中保持可比性。在GreenLight-Gym中,该框架将标量奖励分解为上述条件化分量;适配GreenLight至第二届自主温室挑战的真实气候数据;并在真实温室数据上对相同分量进行评分。

原文摘要 · Abstract (English)

Greenhouse reinforcement learning can test climate-control ideas at a speed and scale that is difficult to achieve with crop experiments alone. For smart-greenhouse control, however, a single simulator return is not enough: a grower or control engineer also needs to know when the policy heats, enriches CO2, vents, manages humidity, deploys screens, or uses lamps.We propose a reproducible calibration-first reward audit framework that keeps named greenhouse-control reward components comparable across simulator training, facility-adapted rollouts, logged Autonomous Greenhouse Challenge records, and actuator-rule distillation. In GreenLight-Gym, the framework decomposes the scalar reward into conditional temperature, CO2, humidity and vapor-pressure-deficit, screen, and actuation-proxy terms; adapts GreenLight to the second Autonomous Greenhouse Challenge logged climate traces; and scores the same components on logged greenhouse data.

强化学习温室控制奖励分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。