arXiv:2505.21201cs.LG2025-05被引 5

用环境与经济数据,帮印度农民选高产高利润作物。

Crop recommendation with machine learning: leveraging environmental and economic factors for optimal crop selection

  • 结合气候、土壤和价格等19类因素建模,用滞后变量处理时间依赖性。
  • 滞后变量法使随机森林模型准确率达83.62%,优于传统交叉验证。
  • 适合农业政策制定者和智慧农业系统开发者参考。

农业是印度粮食生产、经济增长和就业的主渠道,但面临生产力低下、资源压力加剧及气候变化等挑战。尽管推广了绿色革命、灌溉、优质种子和有机农业,成效仍不理想。现有推荐系统多仅关注环境因素和地理区域,难以精准预测高产高利润作物。本研究整合环境与经济因素,基于15个州19种作物数据,采用随机森林(RF)与支持向量机(SVM)模型,通过10折交叉验证、时间序列分割和滞后变量方法进行评估。10折交叉验证表现优异(RF: 99.96%, SVM: 94.71%),但存在过拟合风险;引入时间顺序后性能下降(RF: 78.55%, SVM: 71.18%)。采用滞后变量方法后,模型性能提升(RF: 83.62%, SVM: 74.38%),有效处理时间依赖性,增强对动态农业条件的适应能力。结果表明,在滞后变量框架下,随机森林模型最适用于印度的作物推荐。

原文摘要 · Abstract (English)

Agriculture constitutes a primary source of food production, economic growth and employment in India, but the sector is confronted with low farm productivity and yields aggravated by increased pressure on natural resources and adverse climate change variability. Efforts involving green revolution, land irrigations, improved seeds and organic farming have yielded suboptimal outcomes. The adoption of computational tools like crop recommendation systems offers a new way to provide insights and help farmers tackle low productivity. However, most agricultural recommendation systems in India focus narrowly on environmental factors and regions, limiting accurate predictions of high-yield, profitable crops. This study uses environmental and economic factors with 19 crops across 15 states to develop and evaluate Random Forest and SVM models using 10-fold Cross Validation, Time-series Split, and Lag Variables. The 10-fold cross validation showed high accuracy (RF: 99.96%, SVM: 94.71%) but raised overfitting concerns. Introducing temporal order, better reflecting real-world conditions, reduced performance (RF: 78.55%, SVM: 71.18%) in the Time-series Split.To further increase the model accuracy while maintaining the temporal order, the Lag Variables approach was employed, which resulted in improved performance (RF: 83.62%, SVM: 74.38%) compared to the 10-fold cross validation approach. Overall, the models in the Time-series Split and Lag Variable Approaches offer practical insights by handling temporal dependencies and enhancing its adaptability to changing agricultural conditions over time. Consequently, the study shows the Random Forest model developed based on the Lag Variables as the most preferred algorithm for optimal crop recommendation in the Indian context.

作物推荐机器学习农业决策时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。