arXiv:2511.11880cs.LGcs.AI2025-11

对比Transformer与RNN在森林碳吸收预测中的表现,发现两者精度相近但各有优势。

Transformers vs. Recurrent Models for Estimating Forest Gross Primary Production

  • 用多源数据输入,比较GPT-2(Transformer)和LSTM(RNN)模型的性能
  • LSTM整体更准,但GPT-2在极端事件中表现更好,且需更长时间窗口
  • 辐射是主要影响因子,卫星遥感数据贡献显著,适合生态监测研究者

监测森林二氧化碳吸收(总初级生产力,GPP)的时空动态,仍是陆地生态系统研究的核心挑战。虽然涡流相关(EC)塔提供高频率观测,但空间覆盖有限,难以支持大尺度评估。遥感提供了可扩展的替代方案,但多数方法依赖单传感器光谱指数和统计模型,难以捕捉GPP的复杂时间动态。深度学习与数据融合的进展为更好地表征植被过程的时间特性带来了新机遇,但对先进深度学习模型在多模态GPP预测中的比较评估仍较缺乏。本文探讨了两种代表性模型:1)GPT-2(Transformer架构)和2)长短期记忆网络(LSTM,递归神经网络),使用多变量输入进行GPP预测。总体而言,两者精度相似,但LSTM整体表现更优,而GPT-2在极端事件中更出色。时间上下文长度分析表明,LSTM在显著更短的输入窗口下即可达到相似精度,凸显两类架构在精度与效率间的权衡。特征重要性分析显示,辐射是主导预测因子,其次为哨兵-2(Sentinel-2)、MODIS地表温度和哨兵-1(Sentinel-1)的贡献。结果表明,模型架构、上下文长度与多模态输入共同决定GPP预测性能,为未来陆地碳动态监测的深度学习框架发展提供指导。

原文摘要 · Abstract (English)

Monitoring the spatiotemporal dynamics of forest CO$_2$ uptake (Gross Primary Production, GPP), remains a central challenge in terrestrial ecosystem research. While Eddy Covariance (EC) towers provide high-frequency estimates, their limited spatial coverage constrains large-scale assessments. Remote sensing offers a scalable alternative, yet most approaches rely on single-sensor spectral indices and statistical models that are often unable to capture the complex temporal dynamics of GPP. Recent advances in deep learning (DL) and data fusion offer new opportunities to better represent the temporal dynamics of vegetation processes, but comparative evaluations of state-of-the-art DL models for multimodal GPP prediction remain scarce. Here, we explore the performance of two representative models for predicting GPP: 1) GPT-2, a transformer architecture, and 2) Long Short-Term Memory (LSTM), a recurrent neural network, using multivariate inputs. Overall, both achieve similar accuracy. But, while LSTM performs better overall, GPT-2 excels during extreme events. Analysis of temporal context length further reveals that LSTM attains similar accuracy using substantially shorter input windows than GPT-2, highlighting an accuracy-efficiency trade-off between the two architectures. Feature importance analysis reveals radiation as the dominant predictor, followed by Sentinel-2, MODIS land surface temperature, and Sentinel-1 contributions. Our results demonstrate how model architecture, context length, and multimodal inputs jointly determine performance in GPP prediction, guiding future developments of DL frameworks for monitoring terrestrial carbon dynamics.

GPP预测深度学习遥感碳循环

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。