arXiv:2601.14044cs.CV2026-01

解决气象领域视觉语言模型推理自相矛盾问题,提升决策可信度。

Weather-R1: Logically Consistent Reinforcement Fine-Tuning for Multimodal Reasoning in Meteorology

  • 引入逻辑一致性奖励,修复强化微调中的自相矛盾
  • 在气象推理基准上性能提升9.8个百分点,超越基线与原版Qwen2.5-VL
  • 首个具备逻辑忠实性的气象多模态推理模型,适合高风险场景应用

尽管视觉语言模型(VLM)推理能力不断进步,其在气象领域的应用仍受限于领域差距和推理可信度不足。主流强化微调(RFT)会引发自相矛盾推理(Self-Contra),即模型推理过程与其最终答案矛盾,这在高风险气象决策中不可接受。为此,我们构建了天气问答(WeatherQA)这一新型气象多模态推理基准,并提出逻辑一致强化微调(LoCo-RFT),通过引入逻辑一致性奖励解决自相矛盾问题。此外,我们推出了Weather-R1——目前已知首个在气象领域具备逻辑忠实性的推理型VLM。实验表明,Weather-R1在WeatherQA上的性能比基线高出9.8个百分点,优于监督微调与传统RFT,甚至超越原始Qwen2.5-VL-32B。结果验证了LoCo-RFT的有效性及Weather-R1的优越性。相关基准与代码已开源:https://github.com/Marcowky/Weather-R1。

原文摘要 · Abstract (English)

While Vision Language Models (VLMs) show advancing reasoning capabilities, their application in meteorology is constrained by a domain gap and a reasoning faithfulness gap. Specifically, mainstream Reinforcement Fine-Tuning (RFT) can induce Self-Contradictory Reasoning (Self-Contra), where the model's reasoning contradicts its final answer, which is unacceptable in such a high-stakes domain. To address these challenges, we construct WeatherQA, a novel multimodal reasoning benchmark in meteorology. We also propose Logically Consistent Reinforcement Fine-Tuning (LoCo-RFT), which resolves Self-Contra by introducing a logical consistency reward. Furthermore, we introduce Weather-R1, the first reasoning VLM with logical faithfulness in meteorology, to the best of our knowledge. Experiments demonstrate that Weather-R1 improves performance on WeatherQA by 9.8 percentage points over the baseline, outperforming Supervised Fine-Tuning and RFT, and even surpassing the original Qwen2.5-VL-32B. These results highlight the effectiveness of our LoCo-RFT and the superiority of Weather-R1. Our benchmark and code are available at https://github.com/Marcowky/Weather-R1.

气象推理多模态强化学习逻辑一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。