用多模态模型提升孟加拉语灾害分类准确率,助力灾情实时响应。
BanglaMM-Disaster: A Multimodal Transformer-Based Deep Learning Framework for Multiclass Disaster Classification in Bangla
- 融合文本与图像的Transformer-CNN框架,支持孟加拉语数据
- 83.76%准确率,较纯文本提升3.84%,较纯图像提升16.91%
- 适用于低资源环境下的灾害监测,尤其改善模糊案例分类
自然灾害对孟加拉国仍是重大挑战,因此实时监测与快速响应系统至关重要。本文提出BanglaMM-Disaster,一种基于深度学习的端到端多模态框架,用于孟加拉语社交媒体内容中的灾害分类。我们构建了一个包含5,037条孟加拉语社交帖子的新数据集,每条包含一个标题和对应图像,标注为九类灾害相关类别。模型结合Transformer文本编码器(BanglaBERT、mBERT、XLM-RoBERTa)与CNN主干网络(ResNet50、DenseNet169、MobileNetV2),通过早期融合处理双模态信息。最佳模型达到83.76%准确率,较最优文本基线提升3.84%,较图像基线提升16.91%。分析显示各类别误分类率均下降,尤其在模糊样本上表现显著改善。本工作填补了孟加拉语多模态灾害分析的空白,验证了多源数据融合在低资源环境实时灾情响应中的优势。
原文摘要 · Abstract (English)
Natural disasters remain a major challenge for Bangladesh, so real-time monitoring and quick response systems are essential. In this study, we present BanglaMM-Disaster, an end-to-end deep learning-based multimodal framework for disaster classification in Bangla, using both textual and visual data from social media. We constructed a new dataset of 5,037 Bangla social media posts, each consisting of a caption and a corresponding image, annotated into one of nine disaster-related categories. The proposed model integrates transformer-based text encoders, including BanglaBERT, mBERT, and XLM-RoBERTa, with CNN backbones such as ResNet50, DenseNet169, and MobileNetV2, to process the two modalities. Using early fusion, the best model achieves 83.76% accuracy. This surpasses the best text-only baseline by 3.84% and the image-only baseline by 16.91%. Our analysis also shows reduced misclassification across all classes, with noticeable improvements for ambiguous examples. This work fills a key gap in Bangla multimodal disaster analysis and demonstrates the benefits of combining multiple data types for real-time disaster response in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。