融合土壤影像与营养数据,实现精准作物推荐
AgroSense 2.0: Cross-Modal Transformer Fusion with Geospatial Raster Integration and Interpretable Multi-Task Learning for Precision Crop Recommendation

- 用跨模态Transformer实现图像与表格数据的深度交互
- 引入印度全国土壤栅格数据,提升地理空间建模能力
- 通过可解释性分析揭示不同作物的关键影响因素
精准农业中的作物推荐系统长期面临视觉土壤表征与化学养分分析之间的模态鸿沟:两者通常被当作独立问题处理,融合方式也仅限于后期特征拼接。AgroSense 2.0通过三项架构改进解决此问题。首先,引入覆盖印度的七波段土壤栅格数据(india_soil_7bands.tif),将氮、pH、SOC、黏土、砂粒、粉粒和容重编码为32×32的空间块,这一模态此前未在相关工作中出现。其次,以跨模态Transformer融合模块替代简单的特征拼接,使表格型养分特征通过多头注意力机制关注图像表示,实现更丰富的跨模态依赖建模。第三,采用多任务学习框架,联合优化土壤分类与作物推荐,共享主干网络以增强泛化能力。为提升可解释性,使用TreeSHAP分析表格分支,揭示作物相关的养分敏感性:全球范围内湿度和降雨量影响最显著;具体作物中,水稻受降雨主导,玉米受氮钾影响大,咖啡则由湿度与氮素主导。这些解释既呈现农学上一致的规律,也暴露数据集特有的差异,值得深入研究。整体上,该工作构建了一个更严谨、可解释且地理空间感知更强的精准农业推荐框架。
原文摘要 · Abstract (English)
Crop recommendation systems in precision agriculture have long suffered from a fundamental modality gap: visual soil characterization and chemical nutrient profiling are typically treated as independent inference problems, with fusion often reduced to late-stage feature concatenation. AgroSense~2.0 addresses this limitation through three architectural advances. First, we introduce continental-scale geospatial integration via a seven-band soil raster (\texttt{india\_soil\_7bands.tif}) spanning India, encoding Nitrogen, pH, SOC, Clay, Sand, Silt, and Bulk Density as $32\times32$ spatial patches, a modality entirely absent from prior work. Second, we replace naive feature concatenation with a cross-modal Transformer fusion module, where tabular nutrient features attend over image representations via multi-head attention, enabling richer inter-modal dependency modeling than shallow fusion. Third, we adopt a multi-task objective jointly optimizing soil classification and crop recommendation through a shared backbone, improving generalization via complementary cross-task signal. To enhance interpretability, we apply TreeSHAP to the tabular branch, revealing crop-conditioned nutrient sensitivity: humidity and rainfall emerge as the most influential features globally, while crop-specific profiles diverge meaningfully rainfall dominates rice, nitrogen and potassium dominate maize, and humidity and nitrogen dominate coffee. These explanations provide transparency into model decisions and surface both agronomically consistent patterns and dataset-specific divergences worth further study. Together, these contributions establish AgroSense~2.0 as a more principled, interpretable, and geospatially grounded framework for precision agriculture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。