arXiv:2603.29981cs.LGstat.ML2026-03被引 2

解决空间预测中验证数据偏差问题,让模型评估更贴近真实部署环境。

Aligning Validation with Deployment in Spatial Prediction: Target-Weighted Cross-Validation

  • 用空间任务特征加权验证数据,使评估更匹配实际预测场景。
  • 在德国氮氧化物地图任务中,传统方法高估误差达30%以上。
  • 适合样本分布不均的空间建模,尤其适用于环境监测与遥感应用。

空间环境建模中,机器学习模型常用于从分布不均的观测数据生成地图。标准交叉验证(CV)假设验证数据能代表目标域的预测条件,但实际中因采样偏好或聚集,导致性能和不确定性估计严重偏差。本文提出以部署为导向的加权验证框架,包括重要性加权交叉验证(IWCV)和基于校准的目标加权交叉验证(TWCV),利用环境协变量和预测距离等空间有意义的任务描述符进行加权。模拟实验表明,在真实采样设计下,传统非空间与空间CV存在显著偏差,而加权方法显著降低偏差,当验证任务充分覆盖部署任务空间时效果更优。德国氮氧化物(NO₂)浓度制图案例显示,标准CV因采样偏差可能高估预测误差超过30%,而加权CV结果更符合实际部署条件。该框架将验证任务生成与风险估计分离,为样本分布与预测域不一致的空间预测提供了实用的性能评估改进方法。

原文摘要 · Abstract (English)

Reliable estimation of predictive performance is essential for spatial environmental modeling, where machine-learning models are used to generate maps from unevenly distributed observations. Standard cross-validation (CV) assumes that validation data are representative of prediction conditions across the target domain. In practice, this assumption is often violated due to preferential or clustered sampling, leading to biased performance and uncertainty estimates. We introduce a deployment-oriented validation framework based on weighted CV that aligns validation tasks with the distribution of prediction tasks across a specified domain. The framework includes importance-weighted cross-validation (IWCV) and a calibration-based approach, Target-Weighted Cross-Validation (TWCV), which uses spatially meaningful task descriptors such as environmental covariates and prediction distance. Simulation experiments show that conventional non-spatial and spatial CV strategies can exhibit substantial bias under realistic sampling designs, whereas weighted CV approaches substantially reduce this bias when validation tasks adequately cover the deployment-task space. A case study on mapping nitrogen dioxide (NO$_2$) concentrations across Germany demonstrates that standard CV can overestimate prediction error due to sampling bias, while weighted CV yields estimates more consistent with deployment conditions. The framework separates validation task generation from risk estimation and provides a practical approach for improving performance assessment in spatial prediction settings where sample distributions differ from prediction domains.

空间预测交叉验证环境建模加权评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。