arXiv:2410.13384cs.CV2024-10被引 25

用智能代理实现遥感灾情图像的多任务协同解析,提升救灾决策效率。

RescueADI: Adaptive Disaster Interpretation in Remote Sensing Images with Autonomous Agents

  • 基于大模型的智能代理规划并执行多阶段感知任务
  • 在9类复杂请求上准确率比传统方法高9%
  • 适合应急响应、遥感分析等需要综合判断的场景

当前遥感图像灾情解析方法多聚焦于分割、检测或视觉问答等单一任务,难以应对需融合多种感知手段与专业工具的综合分析需求。为此,本文提出自适应灾情解析(ADI)新任务,通过规划并执行一系列关联性强的解析步骤,实现灾情场景的全面分析。为推动该领域研究,我们构建了新数据集RescueADI,包含4044幅高分辨率遥感图像,涵盖16,949个语义掩码、14,483个目标边界框和13,424条解析请求,覆盖九类挑战性请求类型。同时提出一种基于大语言模型驱动的自主代理方法,可无需人工干预完成计数、面积计算、路径规划等复杂任务。在RescueADI上的实验表明,该方法相比现有VQA方法准确率提升9%,验证了其有效性。数据集将公开共享。

原文摘要 · Abstract (English)

Current methods for disaster scene interpretation in remote sensing images (RSIs) mostly focus on isolated tasks such as segmentation, detection, or visual question-answering (VQA). However, current interpretation methods often fail at tasks that require the combination of multiple perception methods and specialized tools. To fill this gap, this paper introduces Adaptive Disaster Interpretation (ADI), a novel task designed to solve requests by planning and executing multiple sequentially correlative interpretation tasks to provide a comprehensive analysis of disaster scenes. To facilitate research and application in this area, we present a new dataset named RescueADI, which contains high-resolution RSIs with annotations for three connected aspects: planning, perception, and recognition. The dataset includes 4,044 RSIs, 16,949 semantic masks, 14,483 object bounding boxes, and 13,424 interpretation requests across nine challenging request types. Moreover, we propose a new disaster interpretation method employing autonomous agents driven by large language models (LLMs) for task planning and execution, proving its efficacy in handling complex disaster interpretations. The proposed agent-based method solves various complex interpretation requests such as counting, area calculation, and path-finding without human intervention, which traditional single-task approaches cannot handle effectively. Experimental results on RescueADI demonstrate the feasibility of the proposed task and show that our method achieves an accuracy 9% higher than existing VQA methods, highlighting its advantages over conventional disaster interpretation approaches. The dataset will be publicly available.

灾情解析智能代理遥感图像多任务协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。