提升农业强化学习鲁棒性,通过分阶段增强与领域优先噪声注入。
Progressive Generalization Augmentation with Deeply Coupled RND-PPO and Domain-Prioritized Noise Injection for Robust Crop Management Reinforcement Learning
- 分三阶段渐进式增强训练,平衡学习效率与泛化能力。
- 在佛罗里达实现产量提升8.43%、氮肥效率提升16.42%。
- 针对不同农情变量设计优先级噪声,适配实际种植场景。
在基于gym-DSSAT的玉米灌溉任务中,初步实验发现±2℃温度扰动会使纯环境训练的PPO策略经济收益下降11.9%,暴露现有研究对鲁棒性的忽视。本文解决三个关键瓶颈:早期学习效率与后期泛化能力的权衡、探索中内外奖励的简单叠加、以及忽略农情变量敏感差异的均匀噪声注入。提出三项创新:渐进式泛化增强(PGA),包含清洁训练(0-800轮)、渐进增强(800-1200轮)、全增强(1200-2000轮)三阶段课程;深度耦合RND-PPO架构,含双通道GAE归一化、进度衰减内在系数与语义离散化;领域优先噪声注入,支持层级激活。实验显示:在佛罗里达较SOTA BERT-DQN提升8.43%产量、16.42%氮肥效率;在萨拉戈萨产量提升5.61%(但因地中海气候导致经济得分低3.67%);在联合扰动下性能保留率94.4%(对比基线80.0%)。所有实验使用5个随机种子,在NVIDIA A100 GPU上运行,每轮约4.2±0.3小时(2000轮,2048步缓冲区,64小批量)。
原文摘要 · Abstract (English)
Our preliminary experiments on gym-DSSAT maize irrigation tasks revealed that +/-2 degrees C temperature noise causes an 11.9% reduction in economic returns for PPO policies trained under clean conditions - a systematic robustness deficit that existing research has not adequately addressed. This paper tackles three interconnected limitations impeding practical deployment of agricultural RL systems: the trade-off between early-stage learning efficiency and late-stage generalization capability; the naive additive combination of intrinsic and extrinsic rewards in exploration-augmented PPO; and uniform measurement noise injection strategies that disregard empirically validated differential sensitivity across agricultural state variables. We introduce three systematic innovations: Progressive Generalization Augmentation (PGA) implementing a three-phase curriculum (clean training 0-800 episodes, progressive 800-1200, full augmentation 1200-2000); a deeply coupled RND-PPO architecture with dual-channel GAE normalization, progress-decayed intrinsic coefficients, and semantic discretization; and domain-prioritized noise injection with hierarchical activation. Our experimental evaluation demonstrates: 8.43% yield improvement and 16.42% nitrogen use efficiency improvement over SOTA BERT-DQN in Florida; 5.61% yield improvement in Zaragoza (though 3.67% lower economic score due to challenging Mediterranean climate); and 94.4% vs 80.0% performance retention under combined perturbations. All experiments used 5 random seeds on NVIDIA A100 GPUs with 4.2+/-0.3 hours per run (2000 episodes, 2048-step buffer, 64 mini-batch size).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。