arXiv:2510.08043cs.CLcs.LG2025-10被引 1

测试大模型对全球气温数据的记忆能力,发现其在高海拔地区表现差。

Climate Knowledge in Large Language Models

  • 用1°分辨率网格点测试模型记忆特定地点7月平均气温
  • 误差3-6℃,海拔超1500米时误差飙升至5-13℃
  • 加入地理位置信息可降错27%,但无法捕捉温升空间分布

大型语言模型(LLMs)在气候相关应用中日益普及,准确掌握其内部气候知识对可靠性与误信风险评估至关重要。尽管应用广泛,现有模型对参数化气候常识的回忆能力仍缺乏系统评估。本文聚焦典型问题:1991–2020年指定地点的7月地表气温均值,构建全球陆地1°分辨率查询网格,提供坐标与位置描述,并以ERA5再分析数据为基准验证。结果显示,模型编码了非平凡的气候结构,能捕捉纬度与地形模式,均方根误差(RMSE)为3–6℃,偏差±1℃。然而,山区与高纬度地区存在空间一致性的误差。海拔高于1500米时,RMSE达5–13℃,远高于低海拔地区的2–4℃。引入地理上下文(国家、城市、区域)可使误差平均降低27%,且大模型对此更敏感。模型虽能反映1950–1974年与2000–2024年间的全球平均变暖幅度,却无法再现温度变化的空间格局,这直接影响气候变化评估。该结果表明,模型虽能把握当前气候分布,但在表达长期温度变化的区域与局部特征方面存在局限。本研究框架为量化LLMs中的参数化气候知识提供了可复现的基准,补充了现有气候传播评估体系。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed for climate-related applications, where understanding internal climatological knowledge is crucial for reliability and misinformation risk assessment. Despite growing adoption, the capacity of LLMs to recall climate normals from parametric knowledge remains largely uncharacterized. We investigate the capacity of contemporary LLMs to recall climate normals without external retrieval, focusing on a prototypical query: mean July 2-m air temperature 1991-2020 at specified locations. We construct a global grid of queries at 1° resolution land points, providing coordinates and location descriptors, and validate responses against ERA5 reanalysis. Results show that LLMs encode non-trivial climate structure, capturing latitudinal and topographic patterns, with root-mean-square errors of 3-6 °C and biases of $\pm$1 °C. However, spatially coherent errors remain, particularly in mountains and high latitudes. Performance degrades sharply above 1500 m, where RMSE reaches 5-13 °C compared to 2-4 °C at lower elevations. We find that including geographic context (country, city, region) reduces errors by 27% on average, with larger models being most sensitive to location descriptors. While models capture the global mean magnitude of observed warming between 1950-1974 and 2000-2024, they fail to reproduce spatial patterns of temperature change, which directly relate to assessing climate change. This limitation highlights that while LLMs may capture present-day climate distributions, they struggle to represent the regional and local expression of long-term shifts in temperature essential for understanding climate dynamics. Our evaluation framework provides a reproducible benchmark for quantifying parametric climate knowledge in LLMs and complements existing climate communication assessments.

大模型气候知识温度预测地理上下文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。