用多视角图像和视觉语言模型,零样本快速评估野火损毁程度。
Automated Wildfire Damage Assessment from Multi view Ground level Imagery Via Vision Language Models
- 利用预训练多模态大模型,通过多角度图像融合实现零样本损伤分类。
- 多视角分析使中等损毁识别准确率显著提升,单视角效果有限。
- 简单提示词即可达到与复杂推理策略相当的精度,适合灾后快速部署。
野火频发且强度加剧,亟需快速精准的财产损毁评估方法。传统方法耗时长,现代计算机视觉模型通常依赖大量标注数据,难以在灾后即时应用。本研究提出一种新颖的零样本框架,利用预训练多模态大语言模型(MLLMs)对地面影像进行损毁分类。以GPT-4o为主模型,对比Qwen2.5-Vision-Language-32-Billion-Instruct,在加州2025年Eaton与Palisades火灾中测试两种流程:端到端推理(Pipeline A)与视觉引导文本分类的解耦流程(Pipeline B)。研究证明了MLLMs融合多视角信息的有效性:单视角难以识别中等损毁,而多视角分析带来显著改进。通过对比基础零样本提示与结构化思维链、自一致性等先进推理策略,发现简单提示已可达到相近精度。
原文摘要 · Abstract (English)
The escalating intensity and frequency of wildfires demand innovative computational methods for rapid and accurate property damage assessment. Traditional methods are often time-consuming, while modern computer vision approaches typically require extensive labeled datasets, hindering immediate post-disaster deployment. This research introduces a novel, zero-shot framework leveraging pre-trained multimodal large language models (MLLMs) to classify damage from ground-level imagery. Using Generative Pre-trained Transformer 4o (GPT-4o) as the primary model with comparative validation against Qwen2.5-Vision-Language-32-Billion-Instruct (Qwen), the research evaluates two pipelines applied to the 2025 Eaton and Palisades fires in California. These pipelines include an end-to-end inference method (Pipeline A) and a decoupled workflow where visual cues drive text-based classification (Pipeline B). A primary contribution of this study is demonstrating the efficacy of MLLMs in synthesizing information from multiple perspectives. The findings show that while single-view assessments struggle to classify intermediate damage, a multi-view analysis yields dramatic improvements. To explore the impact of prompting methods, the research benchmarked a baseline zero-shot and heuristic approach against advance reasoning strategies (Structured-Chain-of-Thought and Self-Consistency). The results indicate that simple prompting methods achieve a comparable accuracy to the reasoning strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。