arXiv:2604.12831cs.RO2026-04

用视觉语言模型提升火场多智能体协同导航能力

VULCAN: Vision-Language-Model Enhanced Multi-Agent Cooperative Navigation for Indoor Fire-Disaster Response

论文配图:VULCAN: Vision-Language-Model Enhanced Multi-Agent Cooperative Navigation for Indoor Fire-Disaster Response
图 1 · 摘自论文原文
  • 融合视觉与语言模型,实现火场多模态感知
  • 在烟雾扩散等动态环境下仍保持高效探索能力
  • 适合消防机器人、应急搜救系统研究者参考

室内火灾因浓烟、高温及环境动态变化,对自主搜救带来严峻挑战。多智能体协同导航可比单智能体更快更广地探索,但现有方法多基于视觉且针对常规室内场景,在火灾驱动的动态条件下性能显著下降。本文提出VULCAN框架,基于多模态感知与视觉语言模型(VLMs),专为室内火灾救援设计。我们扩展了Habitat-Matterport3D基准,模拟真实物理火灾场景,包括烟雾扩散、热危害与传感器退化。在正常与火灾环境中评估多个代表性多智能体协同导航基线,结果揭示现有方法在火灾场景中的关键失效模式,凸显鲁棒感知与灾祸感知规划的重要性。

原文摘要 · Abstract (English)

Indoor fire disasters pose severe challenges to autonomous search and rescue due to dense smoke, high temperatures, and dynamically evolving indoor environments. In such time-critical scenarios, multi-agent cooperative navigation is particularly useful, as it enables faster and broader exploration than single-agent approaches. However, existing multi-agent navigation systems are primarily vision-based and designed for benign indoor settings, leading to significant performance degradation under fire-driven dynamic conditions. In this paper, we present VULCAN, a multi-agent cooperative navigation framework based on multi-modal perception and vision-language models (VLMs), tailored for indoor fire disaster response. We extend the Habitat-Matterport3D benchmark by simulating physically realistic fire scenarios, including smoke diffusion, thermal hazards, and sensor degradation. We evaluate representative multi-agent cooperative navigation baselines under both normal and fire-driven environments. Our results reveal critical failure modes of existing methods in fire scenarios and underscore the necessity of robust perception and hazard-aware planning for reliable multi-agent search and rescue.

多智能体火场救援视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。