arXiv:2506.00238cs.CVcs.CL2025-06中稿 · the 2025 IEEE Inte…被引 4

零样本视觉问答框架,无需微调即可应对灾后新问题

ZeShot-VQA: Zero-Shot Visual Question Answering Framework with Answer Mapping for Natural Disaster Damage Assessment

  • 基于大模型的零样本学习,直接生成未见答案
  • 在洪水损毁数据集上实现无微调推理
  • 适合快速响应突发灾害的应急决策场景

自然灾害常波及广阔区域并严重破坏基础设施,及时高效的响应对减轻社区影响至关重要,数据驱动方法是首选。视觉问答(VQA)模型有助于管理团队深入理解损毁情况。然而,现有模型多只能从预定义答案列表中选择最佳答案,无法回答训练中未出现的新问题,若需新增问题类型则必须重新收集标注数据并微调模型,耗时费力。近年来,大规模视觉语言模型(VLMs)受到广泛关注,其在海量数据上训练,具备强大的跨模态能力,常可零样本适配下游任务。本文提出一种基于VLM的零样本视觉问答方法(ZeShot-VQA),并在灾后洪水损毁数据集FloodNet上验证性能。该方法无需微调即可应用于新数据集,且能处理并生成训练中未见过的答案,展现出高度灵活性。

原文摘要 · Abstract (English)

Natural disasters usually affect vast areas and devastate infrastructures. Performing a timely and efficient response is crucial to minimize the impact on affected communities, and data-driven approaches are the best choice. Visual question answering (VQA) models help management teams to achieve in-depth understanding of damages. However, recently published models do not possess the ability to answer open-ended questions and only select the best answer among a predefined list of answers. If we want to ask questions with new additional possible answers that do not exist in the predefined list, the model needs to be fin-tuned/retrained on a new collected and annotated dataset, which is a time-consuming procedure. In recent years, large-scale Vision-Language Models (VLMs) have earned significant attention. These models are trained on extensive datasets and demonstrate strong performance on both unimodal and multimodal vision/language downstream tasks, often without the need for fine-tuning. In this paper, we propose a VLM-based zero-shot VQA (ZeShot-VQA) method, and investigate the performance of on post-disaster FloodNet dataset. Since the proposed method takes advantage of zero-shot learning, it can be applied on new datasets without fine-tuning. In addition, ZeShot-VQA is able to process and generate answers that has been not seen during the training procedure, which demonstrates its flexibility.

零样本学习视觉问答灾后评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。