arXiv:2511.17171cs.CVcs.LG2025-11被引 3

用思维链引导视觉模型,跨大陆预测野火风险图

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle

  • 基于多模态数据和思维链推理生成风险图
  • 在美国训练、欧洲测试仍保持高精度表现
  • 首次实现可解释的跨大陆野火风险建模

野火风险预测是需要融合视觉、气候与地理因素的复杂空间推理任务。现有方法缺乏因果推理和多模态理解能力,难以可靠泛化。我们构建了FireScope-Bench——一个大规模数据集与基准,整合美国的哨兵-2影像与气候数据,以及欧洲真实野火事件,用于跨大陆评估。在此基础上,提出FireScope框架:基于视觉语言模型,结合强化学习与视觉监督,生成带有互补推理轨迹的风险栅格图。模型在美国训练后在欧洲测试仍表现优异,专家反馈与自动化分析证实其推理轨迹忠实且语义清晰。结果表明,推理能有效提升栅格预测模型的泛化性与可解释性。据我们所知,这是首个(1)证明语言推理可提升视觉生成泛化性能,(2)实现跨大陆高分辨率野火风险建模,(3)系统研究多模态火灾风险模型跨大陆鲁棒性的框架。数据与代码将公开。

原文摘要 · Abstract (English)

Predicting wildfire risk is a reasoning-intensive spatial problem that requires the integration of visual, climatic, and geographic factors to infer continuous risk maps. Existing methods lack the causal reasoning and multimodal understanding required for reliable generalization. We introduce FireScope-Bench, a large-scale dataset and benchmark that couples Sentinel-2 imagery and climate data with expert-defined risk rasters across the USA, and real wildfire events in Europe for cross-continental evaluation. Building on this dataset, we propose FireScope, a VLM-based reasoning-to-generation framework that learns from both reinforcement learning and visual supervision to predict risk rasters with complementary reasoning traces. When trained in the USA and tested in Europe, FireScope achieves substantial performance gains, while expert feedback and automated analysis confirm that its reasoning traces are faithful and semantically meaningful. Our findings demonstrate that reasoning can ground raster prediction models, improving both generalization and interpretability. To our knowledge, this is the first framework to (1) demonstrate that language-based reasoning can improve generalization in visual generation, (2) propose a high-resolution wildfire risk model that can be applied across continents, and (3) enable systematic studies of robust cross-continental generalization for multimodal fire risk models. We believe that FireScope-Bench has the potential to serve as a foundation for advancing reasoning-driven, interpretable and generalizable spatial modeling. Data and source code will be made publicly available.

野火预测视觉语言模型跨大陆泛化可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。