arXiv:2504.13231cs.CVcs.AI2025-04

构建加拿大野火社交媒体多模态数据集,助力实时灾情研判

WildFireCan-MMD: A Multimodal Dataset for Classification of User-Generated Content During Wildfires in Canada

  • 收集并标注加拿大野火期间推特多模态数据,覆盖12个关键主题
  • 自训练模型达84.48%准确率,显著优于零样本和基线方法
  • 数据集支持灾情趋势分析,强调本地化数据对应急响应的重要性

野火期间快速获取信息至关重要,但传统数据源耗时且成本高。社交媒体提供实时更新,但从中提取有效信息仍具挑战。本文聚焦多模态野火社交媒体数据,尽管现有数据集中已有此类内容,但在加拿大背景下仍严重不足。我们提出 WildFireCan-MMD,一个基于近期加拿大野火的推特多模态数据集,涵盖十二个关键主题的标注。我们在该数据集上评估了零样本视觉语言模型,并与自训练模型及基线分类器进行对比。结果表明,尽管基线方法和零样本提示可快速部署,但在有标注数据时,自训练模型表现更优。最佳自训练模型达到84.48%的f-score,超越视觉语言模型和基线分类器。我们还展示了如何利用该模型分析大规模未标注数据以揭示灾情趋势。本数据集推动未来野火应对研究,研究发现强调定制化数据集和任务特定训练的重要性。尤其重要的是,此类数据集应具有地域针对性,因灾害响应需求随区域与情境而异。

原文摘要 · Abstract (English)

Rapid information access is vital during wildfires, yet traditional data sources are slow and costly. Social media offers real-time updates, but extracting relevant insights remains a challenge. In this work, we focus on multimodal wildfire social media data, which, although existing in current datasets, is currently underrepresented in Canadian contexts. We present WildFireCan-MMD, a new multimodal dataset of X posts from recent Canadian wildfires, annotated across twelve key themes. We evaluate zero-shot vision-language models on this dataset and compare their results with those of custom-trained and baseline classifiers. We show that while baseline methods and zero-shot prompting offer quick deployment, custom-trained models outperform them when labelled data is available. Our best-performing custom model reaches 84.48% f-score, outperforming VLMs and baseline classifiers. We also demonstrate how this model can be used to uncover trends during wildfires, through the collection and analysis of a large unlabeled dataset. Our dataset facilitates future research in wildfire response, and our findings highlight the importance of tailored datasets and task-specific training. Importantly, such datasets should be localized, as disaster response requirements vary across regions and contexts.

多模态数据野火监测社交媒体分析本地化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。