arXiv:2411.01511cs.MAcs.CL2024-11被引 24

用多模型协作自动完成灾后评估与响应,提升效率

Integration of Large Vision Language Models for Efficient Post-disaster Damage Assessment and Reporting

  • 四类专用视觉语言模型协同,由GPT-4统筹决策
  • 实测可显著缩短人类响应时间,优化资源分配
  • 适合应急团队与非专业人士快速参与灾后管理

传统自然灾害应对依赖大量协调工作,速度与效率至关重要。然而人力限制常导致关键行动延迟,加剧人员与经济损失。智能体化大型视觉语言模型(LVLMs)为解决此问题提供新路径,尤其在提升欠发达地区韧性与资源获取方面具有重大社会经济价值。本文提出首个基于多LVLM的灾后管理框架DisasTeller,可自动执行现场评估、紧急预警、资源调配与恢复规划等任务。通过协调四个专业化LVLM智能体,并以GPT-4为核心模型,DisasTeller实现灾难响应活动的自主执行,显著减少人工操作时间并优化资源分布。通过LVLM与人类双重评估验证,该框架在简化响应流程、提升效率方面表现优异,不仅支持专业团队,也使非专业人士能便捷接入灾后管理流程,弥合传统方法与智能驱动效率之间的差距。

原文摘要 · Abstract (English)

Traditional natural disaster response involves significant coordinated teamwork where speed and efficiency are key. Nonetheless, human limitations can delay critical actions and inadvertently increase human and economic losses. Agentic Large Vision Language Models (LVLMs) offer a new avenue to address this challenge, with the potential for substantial socio-economic impact, particularly by improving resilience and resource access in underdeveloped regions. We introduce DisasTeller, the first multi-LVLM-powered framework designed to automate tasks in post-disaster management, including on-site assessment, emergency alerts, resource allocation, and recovery planning. By coordinating four specialised LVLM agents with GPT-4 as the core model, DisasTeller autonomously implements disaster response activities, reducing human execution time and optimising resource distribution. Our evaluations through both LVLMs and humans demonstrate DisasTeller's effectiveness in streamlining disaster response. This framework not only supports expert teams but also simplifies access to disaster management processes for non-experts, bridging the gap between traditional response methods and LVLM-driven efficiency.

灾后评估多智能体视觉语言模型自动化响应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。