arXiv:2506.10419cs.LG2025-06

用聚类优化采样点分布,提升土壤碳监测代表性。

Data-Driven Soil Organic Carbon Sampling: Integrating Spectral Clustering with Conditioned Latin Hypercube Optimization

  • 先聚类分区再优化采样,确保小环境区也被覆盖。
  • 相比传统方法,采样点在环境特征空间更均匀分布。
  • 适合需要高精度土壤碳建模的研究者使用。

土壤有机碳(SOC)监测常依赖环境协变量选择代表性采样点。本文提出一种新型混合方法:将谱聚类(无监督机器学习)与条件拉丁超立方采样(cLHS)结合,以提升采样代表性。该方法首先利用多变量协变量数据,通过谱聚类将研究区划分为K个同质区域;随后在每个区域内应用cLHS,选取能全面捕捉环境条件多样性的采样点。该混合谱聚类-cLHS方法可有效避免传统cLHS忽略重要但小规模环境簇的问题。在真实SOC制图数据集上的实验表明,该方法在协变量特征空间和空间异质性上均实现更均匀的覆盖,优于标准cLHS。这种改进的采样设计有助于为机器学习模型提供更均衡的训练数据,从而提升土壤有机碳预测的准确性。

原文摘要 · Abstract (English)

Soil organic carbon (SOC) monitoring often relies on selecting representative field sampling locations based on environmental covariates. We propose a novel hybrid methodology that integrates spectral clustering - an unsupervised machine learning technique with conditioned Latin hypercube sampling (cLHS) to enhance the representativeness of SOC sampling. In our approach, spectral clustering partitions the study area into $K$ homogeneous zones using multivariate covariate data, and cLHS is then applied within each zone to select sampling locations that collectively capture the full diversity of environmental conditions. This hybrid spectral-cLHS method ensures that even minor but important environmental clusters are sampled, addressing a key limitation of vanilla cLHS which can overlook such areas. We demonstrate on a real SOC mapping dataset that spectral-cLHS provides more uniform coverage of covariate feature space and spatial heterogeneity than standard cLHS. This improved sampling design has the potential to yield more accurate SOC predictions by providing better-balanced training data for machine learning models.

土壤碳采样优化聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。