通过因果引导的对抗解耦,提升危机分类在未知灾害下的泛化能力。
CAMO: Causality-Guided Adversarial Multimodal Domain Generalization for Crisis Classification
- 利用对抗解耦分离因果与虚假特征,增强模型鲁棒性。
- 在未见过的灾害场景中准确率显著提升,优于现有方法。
- 适用于社交媒体多模态灾情信息快速识别,适合应急响应场景。
社交媒体中的危机分类旨在从多模态帖子中提取可行动的灾害相关信息,对提升态势感知和及时应急响应至关重要。然而,灾害类型差异大,跨未见灾害实现良好泛化仍是长期挑战。现有方法主要依赖深度学习融合文本与视觉线索,在域内设置下表现尚可,但在跨域时性能下降,原因在于:1. 未能解耦虚假与因果特征,导致域偏移下性能退化;2. 无法对齐异构模态表示,阻碍单模态域泛化技术向多模态迁移。为此,本文提出一种因果引导的多模态域泛化(CAMO)框架,结合对抗解耦与统一表征学习。对抗目标促使模型聚焦于域不变的因果特征,实现基于稳定因果机制的泛化分类;统一表征将不同模态特征映射至共享潜在空间,使已有单模态域泛化策略可无缝扩展至多模态场景。在多个数据集上的实验表明,该方法在未见灾害场景中表现最优。
原文摘要 · Abstract (English)
Crisis classification in social media aims to extract actionable disaster-related information from multimodal posts, which is a crucial task for enhancing situational awareness and facilitating timely emergency responses. However, the wide variation in crisis types makes achieving generalizable performance across unseen disasters a persistent challenge. Existing approaches primarily leverage deep learning to fuse textual and visual cues for crisis classification, achieving numerically plausible results under in-domain settings. However, they exhibit poor generalization across unseen crisis types because they 1. do not disentangle spurious and causal features, resulting in performance degradation under domain shift, and 2. fail to align heterogeneous modality representations within a shared space, which hinders the direct adaptation of established single-modality domain generalization (DG) techniques to the multimodal setting. To address these issues, we introduce a causality-guided multimodal domain generalization (MMDG) framework that combines adversarial disentanglement with unified representation learning for crisis classification. The adversarial objective encourages the model to disentangle and focus on domain-invariant causal features, leading to more generalizable classifications grounded in stable causal mechanisms. The unified representation aligns features from different modalities within a shared latent space, enabling single-modality DG strategies to be seamlessly extended to multimodal learning. Experiments on the different datasets demonstrate that our approach achieves the best performance in unseen disaster scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。