arXiv:2511.19557cs.CVcs.AI2025-11被引 1

用两阶段推理框架提升灾后无人机影像的问答准确性与可解释性

Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment

  • 先生成思维链推理路径,再筛选最优答案,提升逻辑连贯性
  • 在洪水和飓风数据集上准确率超现有方法,且无需重新训练
  • 适合应急响应、灾害评估等需要快速决策的场景

灾后及时准确的损毁评估对高效应急响应至关重要。现有基于AI的框架虽能分析无人机获取的航拍图像,但标注数据成本高,数据集规模小且多样性不足。多数方法依赖固定答案空间的分类模型,难以提供新信息。利用基于上下文学习的预训练生成模型虽可实现开放回答,却常产生幻觉或通用回复。为此,我们提出ThiFAN-VQA,一种两阶段视觉问答框架。第一阶段通过思维链提示与上下文学习生成结构化推理路径,实现弱监督下的可解释推理;第二阶段通过答案选择模块评估并选取最一致、最符合语境的答案,显著提升性能。结合定制信息检索系统、领域特定提示与推理引导的选择机制,ThiFAN-VQA融合了零样本与有监督方法的优势。在FloodNet和RescueNet-VQA两个基于无人机的洪水与飓风灾区数据集上的实验表明,该框架在真实灾后评估任务中实现了更高的准确率、可解释性与适应性。

原文摘要 · Abstract (English)

Timely and accurate assessment of damages following natural disasters is essential for effective emergency response and recovery. Recent AI-based frameworks have been developed to analyze large volumes of aerial imagery collected by Unmanned Aerial Vehicles, providing actionable insights rapidly. However, creating and annotating data for training these models is costly and time-consuming, resulting in datasets that are limited in size and diversity. Furthermore, most existing approaches rely on traditional classification-based frameworks with fixed answer spaces, restricting their ability to provide new information without additional data collection or model retraining. Using pre-trained generative models built on in-context learning (ICL) allows for flexible and open-ended answer spaces. However, these models often generate hallucinated outputs or produce generic responses that lack domain-specific relevance. To address these limitations, we propose ThiFAN-VQA, a two-stage reasoning-based framework for visual question answering (VQA) in disaster scenarios. ThiFAN-VQA first generates structured reasoning traces using chain-of-thought (CoT) prompting and ICL to enable interpretable reasoning under limited supervision. A subsequent answer selection module evaluates the generated responses and assigns the most coherent and contextually accurate answer, effectively improve the model performance. By integrating a custom information retrieval system, domain-specific prompting, and reasoning-guided answer selection, ThiFAN-VQA bridges the gap between zero-shot and supervised methods, combining flexibility with consistency. Experiments on FloodNet and RescueNet-VQA, UAV-based datasets from flood- and hurricane-affected regions, demonstrate that ThiFAN-VQA achieves superior accuracy, interpretability, and adaptability for real-world post-disaster damage assessment tasks.

灾后评估视觉问答思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。