arXiv:2502.20139cs.LG2025-02被引 7

31个土壤数据集公开,可用来测试机器学习模型在土壤预测中的表现。

LimeSoDa: A Dataset Collection for Benchmarking of Machine Learning Regressors in Digital Soil Mapping

  • 构建31个全球土壤数据集,含有机质、黏粒和pH等目标变量。
  • 对比4种算法发现:高维光谱数据下线性与支持向量机更优。
  • 特征少于20个时,梯度提升和随机森林表现更好,适合不同场景。

数字土壤制图(DSM)依赖多种统计方法,但针对特定情境选择最优方法仍具挑战。现有研究多基于单一受限数据集,结论可能不完整或误导。为此,我们推出开源数据集集合LimeSoDa,包含来自多个地区的31个田块及农场尺度数据集。每个数据集涵盖三种土壤属性:土壤有机质/碳、黏粒含量和pH,以及由光学、近地和遥感技术获取的特征。所有数据均统一为表格格式,可直接用于建模。通过在全部数据集上比较四种算法——多元线性回归(MLR)、支持向量回归(SVR)、类别提升(CatBoost)和随机森林(RF)——的预测性能,结果显示:无单一算法始终最优,但特定情境下表现差异显著。在高维光谱数据中,MLR与SVR表现更佳,可能因其与主成分分析兼容性好;而在特征数少于20的数据集中,CatBoost与RF显著更优。这表明模型表现高度依赖具体情境,LimeSoDa为提升DSM中统计方法的研发与评估提供了重要资源。

原文摘要 · Abstract (English)

Digital soil mapping (DSM) relies on a broad pool of statistical methods, yet determining the optimal method for a given context remains challenging and contentious. Benchmarking studies on multiple datasets are needed to reveal strengths and limitations of commonly used methods. Existing DSM studies usually rely on a single dataset with restricted access, leading to incomplete and potentially misleading conclusions. To address these issues, we introduce an open-access dataset collection called Precision Liming Soil Datasets (LimeSoDa). LimeSoDa consists of 31 field- and farm-scale datasets from various countries. Each dataset has three target soil properties: (1) soil organic matter or soil organic carbon, (2) clay content and (3) pH, alongside a set of features. Features are dataset-specific and were obtained by optical spectroscopy, proximal- and remote soil sensing. All datasets were aligned to a tabular format and are ready-to-use for modeling. We demonstrated the use of LimeSoDa for benchmarking by comparing the predictive performance of four learning algorithms across all datasets. This comparison included multiple linear regression (MLR), support vector regression (SVR), categorical boosting (CatBoost) and random forest (RF). The results showed that although no single algorithm was universally superior, certain algorithms performed better in specific contexts. MLR and SVR performed better on high-dimensional spectral datasets, likely due to better compatibility with principal components. In contrast, CatBoost and RF exhibited considerably better performances when applied to datasets with a moderate number (< 20) of features. These benchmarking results illustrate that the performance of a method is highly context-dependent. LimeSoDa therefore provides an important resource for improving the development and evaluation of statistical methods in DSM.

土壤制图机器学习数据集基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。