比较三种强化学习方法,提升农民收益与分配公平性。
Comparative Analysis of Multi-Agent Reinforcement Learning Policies for Crop Planning Decision Support
- 采用三种多智能体强化学习策略优化作物规划决策。
- 联合优化策略收益最高,但计算开销大;顺序优化平衡效率与公平。
- 适合关注农业决策公平性与可扩展性的研究者参考。
在印度,多数农民属于小规模或边缘农户,其生计极易受市场饱和和气候风险影响。有效的作物规划可显著提升预期收入,但现有决策支持系统(DSS)常提供通用建议,难以反映实时市场动态及多农户间的互动。本文评估了三种多智能体强化学习(MARL)方法在优化总农户收入和促进公平分配方面的可行性:独立Q学习(IQL),各农户独立决策;逐个优化(ABA),按序调整各农户策略以适应他人;多智能体滚动策略(Multi-agent Rollout),联合优化所有农户行动以实现全局奖励最大化。结果表明,虽然IQL计算效率高(线性运行时间),但协作差,总收益低且收入分配不均;而多智能体滚动策略虽获最高总收益并促进收入公平,但计算资源需求大,难以适用于大规模农户群体;ABA在运行效率与收益优化间取得平衡,具备合理总收益、可接受公平性与可扩展性。研究强调,选择合适的MARL方法对构建个性化、公平的作物规划推荐系统至关重要,推动更适应农户需求的农业决策系统发展。
原文摘要 · Abstract (English)
In India, the majority of farmers are classified as small or marginal, making their livelihoods particularly vulnerable to economic losses due to market saturation and climate risks. Effective crop planning can significantly impact their expected income, yet existing decision support systems (DSS) often provide generic recommendations that fail to account for real-time market dynamics and the interactions among multiple farmers. In this paper, we evaluate the viability of three multi-agent reinforcement learning (MARL) approaches for optimizing total farmer income and promoting fairness in crop planning: Independent Q-Learning (IQL), where each farmer acts independently without coordination, Agent-by-Agent (ABA), which sequentially optimizes each farmer's policy in relation to the others, and the Multi-agent Rollout Policy, which jointly optimizes all farmers' actions for global reward maximization. Our results demonstrate that while IQL offers computational efficiency with linear runtime, it struggles with coordination among agents, leading to lower total rewards and an unequal distribution of income. Conversely, the Multi-agent Rollout policy achieves the highest total rewards and promotes equitable income distribution among farmers but requires significantly more computational resources, making it less practical for large numbers of agents. ABA strikes a balance between runtime efficiency and reward optimization, offering reasonable total rewards with acceptable fairness and scalability. These findings highlight the importance of selecting appropriate MARL approaches in DSS to provide personalized and equitable crop planning recommendations, advancing the development of more adaptive and farmer-centric agricultural decision-making systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。