arXiv:2603.18611cs.CLcs.CV2026-03中稿 · WWW 2026

跨模态迁移推理,让危机社交媒体图文分类更透明可解释。

Cross-Modal Rationale Transfer for Explainable Humanitarian Classification on Social Media

  • 用图文联合模型先提取文本理由,再映射生成图像理由。
  • 在CrisisMMD上提升宏平均F1 2-35%,图像理由准确率高12%。
  • 零样本迁移准确率达80%,适合救灾等实时决策场景。

社交媒体在危机中提供实时信息,涵盖基础设施损毁、人员失踪等多种类别。现有方法对文本和图像进行人道主义分类,但决策过程不透明,限制实际应用。近期工作尝试通过提取推文中的文本理由来增强可解释性,但多局限于文本,缺乏对危机图像的解释。本文提出一种可解释设计的多模态分类框架:首先利用视觉语言变换器学习图文联合表示并提取文本理由;随后通过与文本理由的映射关系,提取图像理由。该方法实现了跨模态理由迁移,显著减少标注成本。基于提取理由进行分类,在CrisisMMD基准数据集上的实验表明,本方法使分类宏平均F1提升2-35%,且能准确提取文本词元与图像块作为理由。人工评估进一步验证,本方法在识别人道主义类别时,图像理由块的准确性提升12%。此外,该方法在零样本模式下适应新数据集,达到80%的准确率。

原文摘要 · Abstract (English)

Advances in social media data dissemination enable the provision of real-time information during a crisis. The information comes from different classes, such as infrastructure damages, persons missing or stranded in the affected zone, etc. Existing methods attempted to classify text and images into various humanitarian categories, but their decision-making process remains largely opaque, which affects their deployment in real-life applications. Recent work has sought to improve transparency by extracting textual rationales from tweets to explain predicted classes. However, such explainable classification methods have mostly focused on text, rather than crisis-related images. In this paper, we propose an interpretable-by-design multimodal classification framework. Our method first learns the joint representation of text and image using a visual language transformer model and extracts text rationales. Next, it extracts the image rationales via the mapping with text rationales. Our approach demonstrates how to learn rationales in one modality from another through cross-modal rationale transfer, which saves annotation effort. Finally, tweets are classified based on extracted rationales. Experiments are conducted over CrisisMMD benchmark dataset, and results show that our proposed method boosts the classification Macro-F1 by 2-35% while extracting accurate text tokens and image patches as rationales. Human evaluation also supports the claim that our proposed method is able to retrieve better image rationale patches (12%) that help to identify humanitarian classes. Our method adapts well to new, unseen datasets in zero-shot mode, achieving an accuracy of 80%.

多模态可解释危机分类跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。