arXiv:2510.26099cs.LGcs.AI2025-10被引 2

用分层评估法揭示天气模型在地球各地的表现差异

SAFE: A Novel Approach to AI Weather Evaluation through Stratified Assessments of Forecasts over Earth

  • 按国家、收入、地貌等维度分层评估模型预测表现
  • 发现所有主流AI天气模型在不同地区精度不一
  • 开源工具支持公平性分析,适合气候公平研究者

机器学习普遍以测试集平均损失评估模型性能,这在气象气候领域意味着对全球地理区域的平均评价,忽略了人类发展和地理分布的非均匀性。本文提出分层天气预测评估框架(SAFE),整合多源数据,按地理网格点的国家、全球子区域、收入水平和地表覆盖类型(陆地或水域)进行分层。该方法可分析模型在每个分层中的表现(如每个国家的预测准确率)。通过应用SAFE评估一系列前沿AI天气模型,发现所有模型在各属性上均存在预测能力差异。基于此构建了不同预报时效和气候变量下的模型公平性基准。该工作首次系统追问:模型在哪些地区表现最好或最差?哪个模型最公平?SAFE工具包已开源,供后续研究使用。

原文摘要 · Abstract (English)

The dominant paradigm in machine learning is to assess model performance based on average loss across all samples in some test set. This amounts to averaging performance geospatially across the Earth in weather and climate settings, failing to account for the non-uniform distribution of human development and geography. We introduce Stratified Assessments of Forecasts over Earth (SAFE), a package for elucidating the stratified performance of a set of predictions made over Earth. SAFE integrates various data domains to stratify by different attributes associated with geospatial gridpoints: territory (usually country), global subregion, income, and landcover (land or water). This allows us to examine the performance of models for each individual stratum of the different attributes (e.g., the accuracy in every individual country). To demonstrate its importance, we utilize SAFE to benchmark a zoo of state-of-the-art AI-based weather prediction models, finding that they all exhibit disparities in forecasting skill across every attribute. We use this to seed a benchmark of model forecast fairness through stratification at different lead times for various climatic variables. By moving beyond globally-averaged metrics, we for the first time ask: where do models perform best or worst, and which models are most fair? To support further work in this direction, the SAFE package is open source and available at https://github.com/N-Masi/safe

天气预测公平性评估分层分析AI气象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。