用多模态模型自动解析地下矿难现场,生成清晰描述提升救援效率。
Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters
- 融合视觉与文本特征,即使在黑暗粉尘中也能准确对齐信息
- 在真实矿难数据集上,生成描述准确率显著高于现有模型
- 适合应急救援、矿山安全系统开发人员使用
地下矿难产生浓雾、尘土和坍塌,严重遮蔽视线,导致人类和传统系统难以掌握现场态势。为此,我们提出MDSE——一种新型多模态视觉-语言框架,可自动生成灾后地下场景的详细文本解释。该框架有三项创新:(i) 上下文感知交叉注意力,在极端退化条件下仍能稳健对齐视觉与文本特征;(ii) 分割感知双路径视觉编码,融合全局与区域特异性嵌入;(iii) 资源高效的基于Transformer的语言模型,在极低计算开销下实现丰富文本生成。为支持该任务,我们构建了首个真实地下灾后场景图像-标题语料库——地下矿难数据集(UMD),可用于严格训练与评估。在UMD及关联基准上的大量实验表明,MDSE显著优于现有最佳图像字幕模型,生成的描述更准确、更具上下文相关性,能捕捉被遮蔽环境中的关键细节,有效提升地下应急响应的情境感知能力。代码已开源:https://github.com/mizanJewel/Multimodal-Disaster-Situation-Explainer。
原文摘要 · Abstract (English)
Underground mining disasters produce pervasive darkness, dust, and collapses that obscure vision and make situational awareness difficult for humans and conventional systems. To address this, we propose MDSE, Multimodal Disaster Situation Explainer, a novel vision-language framework that automatically generates detailed textual explanations of post-disaster underground scenes. MDSE has three-fold innovations: (i) Context-Aware Cross-Attention for robust alignment of visual and textual features even under severe degradation; (ii) Segmentation-aware dual pathway visual encoding that fuses global and region-specific embeddings; and (iii) Resource-Efficient Transformer-Based Language Model for expressive caption generation with minimal compute cost. To support this task, we present the Underground Mine Disaster (UMD) dataset--the first image-caption corpus of real underground disaster scenes--enabling rigorous training and evaluation. Extensive experiments on UMD and related benchmarks show that MDSE substantially outperforms state-of-the-art captioning models, producing more accurate and contextually relevant descriptions that capture crucial details in obscured environments, improving situational awareness for underground emergency response. The code is at https://github.com/mizanJewel/Multimodal-Disaster-Situation-Explainer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。