用Transformer自注意力机制解释多模态作物产量预测,让模型决策更透明。
Intrinsic Explainability of Multimodal Learning for Crop Yield Prediction
- 基于自注意力机制,用Attention Rollout和通用注意力评估特征重要性。
- Transformer模型在子地块和地块层面的预测准确率分别高出0.10和0.04。
- 结果可结合农学知识解读生育期阶段,适合农业AI研究者参考。
多模态学习使机器学习任务能融合多种数据源,有效模拟现实世界中不同因素的交互,尤其在农业领域表现突出。尽管数据模态异质性常需复杂架构,模型可解释性却常被忽视。本研究利用基于Transformer模型的内在可解释性,聚焦于亚地块级别的作物产量预测任务。所用大数据集涵盖多种作物、地区与年份,包含四种输入模态:多光谱卫星影像与气象时间序列、地形高程图及土壤属性。基于自注意力机制,采用Attention Rollout(AR)与Generic Attention(GA)两种方法估计特征归因,并与基于Shapley值的模型无关方法Shapley Value Sampling(SVS)进行对比。此外,提出加权模态激活(WMA)方法评估模态归因,并与SVS结果比较。结果表明,Transformer模型优于卷积与循环网络,在亚地块和地块层面分别取得更高的R²值,提升0.10和0.04。AR在时间维度归因上表现更稳健可靠,经定性与定量验证优于GA与SVS。结合作物生育期信息,可基于已有农学知识解释结果。同时,两种方法在模态归因模式上呈现差异。
原文摘要 · Abstract (English)
Multimodal learning enables various machine learning tasks to benefit from diverse data sources, effectively mimicking the interplay of different factors in real-world applications, particularly in agriculture. While the heterogeneous nature of involved data modalities may necessitate the design of complex architectures, the model interpretability is often overlooked. In this study, we leverage the intrinsic explainability of Transformer-based models to explain multimodal learning networks, focusing on the task of crop yield prediction at the subfield level. The large datasets used cover various crops, regions, and years, and include four different input modalities: multispectral satellite and weather time series, terrain elevation maps and soil properties. Based on the self-attention mechanism, we estimate feature attributions using two methods, namely the Attention Rollout (AR) and Generic Attention (GA), and evaluate their performance against Shapley-based model-agnostic estimations, Shapley Value Sampling (SVS). Additionally, we propose the Weighted Modality Activation (WMA) method to assess modality attributions and compare it with SVS attributions. Our findings indicate that Transformer-based models outperform other architectures, specifically convolutional and recurrent networks, achieving R2 scores that are higher by 0.10 and 0.04 at the subfield and field levels, respectively. AR is shown to provide more robust and reliable temporal attributions, as confirmed through qualitative and quantitative evaluation, compared to GA and SVS values. Information about crop phenology stages was leveraged to interpret the explanation results in the light of established agronomic knowledge. Furthermore, modality attributions revealed varying patterns across the two methods compared.[...]
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。