arXiv:2511.00004cs.CYcs.AI2025-11中稿 · 2025 IEEE 21st Int…

用增强技术提升灾难评估中数据不足问题的模型表现

Multimodal Learning with Augmentation Techniques for Natural Disaster Assessment

  • 采用扩散模型与文本重写增强视觉和文本数据
  • 多模态与多视角学习下小类准确率显著提升
  • 适合灾害监测、应急响应等实际应用场景

自然灾害评估依赖快速准确的信息获取,社交媒体成为重要的实时数据源。然而,现有数据集存在类别不平衡和样本有限的问题,制约了模型的有效训练。本文针对CrisisMMD多模态数据集,探索多种增强方法:视觉数据采用基于扩散的Real Guidance和DiffuseMix;文本数据则使用回译、Transformer改写及图像描述生成增强。在单模态、多模态和多视角学习设置下进行评估。结果表明,所选增强策略能有效提升分类性能,尤其对少数类别改善明显;多视角学习虽具潜力,但仍需优化。本研究为构建更鲁棒的灾难评估系统提供了有效的增强策略。

原文摘要 · Abstract (English)

Natural disaster assessment relies on accurate and rapid access to information, with social media emerging as a valuable real-time source. However, existing datasets suffer from class imbalance and limited samples, making effective model development a challenging task. This paper explores augmentation techniques to address these issues on the CrisisMMD multimodal dataset. For visual data, we apply diffusion-based methods, namely Real Guidance and DiffuseMix. For text data, we explore back-translation, paraphrasing with transformers, and image caption-based augmentation. We evaluated these across unimodal, multimodal, and multi-view learning setups. Results show that selected augmentations improve classification performance, particularly for underrepresented classes, while multi-view learning introduces potential but requires further refinement. This study highlights effective augmentation strategies for building more robust disaster assessment systems.

灾难评估多模态数据增强扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。