用强化学习训练AI写专业气象预报,提升准确性和专业性
AFDBench: A Reasoning-First AI Scientist for NationalWeather Service Forecast Discussions

- 通过结构化天气数据推理生成预报文本
- 强化学习使专业风格匹配度提升至0.619,数据忠实度达0.940
- 适合气象AI研发与高精度预报系统构建者
大语言模型在生成高风险气象文本时会幻觉数值,威胁天气通信安全。我们提出AFDBench,首个评估生成式气象推理的基准,包含13个国家气象局办公室的7,732份专家撰写预报讨论,搭配来自Google WeatherNext 2的真实AI天气输入数据。引入三项互补指标:Met-Align(数值准确性)、Style-Align(专业语体契合度)和Input-Grounding(对源数据的忠实度)。零样本测试显示,开源LLM的Style-Align仅约0.33,输入数据忠实度约0.88,无法写出专业气象员风格或正确使用输入数据。采用领域专用奖励函数的群组相对策略优化(GRPO),针对温度准确性、天气形势合理性与格式合规性进行训练。在两个未见气象局的1,033个样本上,GRPO将Style-Align从0.318提升至0.619,输入数据忠实度从0.881提升至0.940,证明70亿参数模型可通过强化学习学会像专业气象员一样写作并准确解读AI天气数据。
原文摘要 · Abstract (English)
Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather communication. We present AFDBench, an AI meteorologist that generates professional Area Forecast Discussions (AFDs) by reasoning through structured AI weather forecast data from Google's WeatherNext 2. We introduce AFDBench, the first benchmark for evaluating generative meteorological reasoning, comprising 7,732 expert written discussions from 13 National Weather Service (NWS) offices paired with real AI weather forecast inputs, and three complementary metrics: Met-Align (numerical accuracy), Style-Align (professional dialect adherence), and Input-Grounding (fidelity to source weather data). Zero-shot evaluations reveal that open-source LLMs achieve low Style-Align (~0.33) and moderate Input-Grounding (~0.88), failing to write in the professional NWS register or faithfully use their input data. We apply Group Relative Policy Optimization (GRPO) with domain-specific rewards targeting temperature accuracy, synoptic correctness, and format compliance. On 1,033 held-out samples from two unseen NWS offices, GRPO nearly doubles Style-Align from 0.318 to 0.619 and improves Input-Grounding from 0.881 to 0.940, demonstrating that reinforcement learning teaches a 7B-parameter model to write like a professional meteorologist and faithfully interpret AI weather data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。