用表格大模型预测南非作物产量,速度快且省去繁琐准备。
From Rows to Yields: How Foundation Models for Tabular Data Simplify Crop Yield Prediction
- 用表格大模型TabPFN处理遥感与气象数据做产量预测
- 准确率与传统机器学习相当,但训练时间大幅缩短
- 适合需要快速部署的农业监测场景
本文将针对中小规模表格数据的通用模型TabPFN应用于南非次国家级作物产量预测任务。基于23年、覆盖最多8个省份的作物产量数据,利用每10天一次的地球观测(FAPAR和土壤湿度)及网格化气象数据(气温、降水、辐射),按区域提取并月度聚合特征输入模型。在留一年外交叉验证设置下,对比了TabPFN与六种机器学习模型及三种基准模型。结果表明,TabPFN与主流机器学习模型表现相近,显著优于基线模型;更重要的是,其调参时间更短、对特征工程依赖更低,展现出更强的实际应用优势,特别适合对效率和易部署性要求高的真实世界产量预测场景。
原文摘要 · Abstract (English)
We present an application of a foundation model for small- to medium-sized tabular data (TabPFN), to sub-national yield forecasting task in South Africa. TabPFN has recently demonstrated superior performance compared to traditional machine learning (ML) models in various regression and classification tasks. We used the dekadal (10-days) time series of Earth Observation (EO; FAPAR and soil moisture) and gridded weather data (air temperature, precipitation and radiation) to forecast the yield of summer crops at the sub-national level. The crop yield data was available for 23 years and for up to 8 provinces. Covariate variables for TabPFN (i.e., EO and weather) were extracted by region and aggregated at a monthly scale. We benchmarked the results of the TabPFN against six ML models and three baseline models. Leave-one-year-out cross-validation experiment setting was used in order to ensure the assessment of the models capacity to forecast an unseen year. Results showed that TabPFN and ML models exhibit comparable accuracy, outperforming the baselines. Nonetheless, TabPFN demonstrated superior practical utility due to its significantly faster tuning time and reduced requirement for feature engineering. This renders TabPFN a more viable option for real-world operation yield forecasting applications, where efficiency and ease of implementation are paramount.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。