arXiv:2510.03349cs.LGcs.AI2025-10被引 1

用多模态大模型做龙卷风预报,检验LLM在真实复杂任务中的推理能力。

AgentCaster: Reasoning-Guided Tornado Forecasting

  • 端到端使用多模态大模型分析高分辨率气象数据
  • 40天历史数据验证,500+龙卷风报告,12-36小时预报窗口
  • 发现模型易幻觉、误判位置,人类专家仍更优

为评估大语言模型(LLMs)在复杂高影响现实任务中的真实推理能力,我们提出AgentCaster——一个无污染的端到端多模态框架,用于长期高难度的龙卷风预报。该框架处理来自高分辨率对流允许预报档案的异构时空数据。在涵盖多个重大龙卷风爆发事件的40天历史数据上进行评估,包含超过500次龙卷风报告。每日模型从3,625张预报图和40,125个探空数据中交互查询,预报时长达12-36小时。概率性龙卷风风险多边形预测通过投影坐标空间中不重叠风险带的几何比对进行验证。我们提出领域特定的TornadoBench与TornadoHallucination指标,其中TornadoBench对当前最先进模型及领域专家均构成巨大挑战。值得注意的是,人类专家显著优于现有模型,而模型普遍存在幻觉、风险强度过度预测、地理定位不准及复杂动态系统中时空推理能力差等问题。AgentCaster旨在推动提升大模型在关键领域的复杂推理任务表现。

原文摘要 · Abstract (English)

There is a growing need to evaluate Large Language Models (LLMs) on complex, high-impact, real-world tasks to assess their true readiness as reasoning agents. To address this gap, we introduce AgentCaster, a contamination-free framework employing multimodal LLMs end-to-end for the challenging, long-horizon task of tornado forecasting. Within AgentCaster, models interpret heterogeneous spatiotemporal data from a high-resolution convection-allowing forecast archive. We assess model performance over a 40-day period featuring diverse historical data, spanning several major tornado outbreaks and including over 500 tornado reports. Each day, models query interactively from a pool of 3,625 forecast maps and 40,125 forecast soundings for a forecast horizon of 12-36 hours. Probabilistic tornado-risk polygon predictions are verified against ground truths derived from geometric comparisons across disjoint risk bands in projected coordinate space. To quantify accuracy, we propose domain-specific TornadoBench and TornadoHallucination metrics, with TornadoBench highly challenging for both LLMs and domain expert human forecasters. Notably, human experts significantly outperform state-of-the-art models, which demonstrate a strong tendency to hallucinate and overpredict risk intensity, struggle with precise geographic placement, and exhibit poor spatiotemporal reasoning in complex, dynamically evolving systems. AgentCaster aims to advance research on improving LLM agents for challenging reasoning tasks in critical domains.

大模型推理龙卷风预报多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。