通过分层跨模态不一致建模,提升社交媒体讽刺与网络欺凌识别准确率。
HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

- 构建分词、短语、全局三层次的跨模态不一致图网络,捕捉多粒度语义矛盾。
- 在MMSD和MultiBully数据集上分别达到85.74%准确率和68.66%宏F1,优于现有方法。
- 适合研究多模态情感分析、社交文本安全检测的开发者与研究人员参考。
多模态讽刺与网络欺凌检测仍具挑战性,因真实语义常源于文本与视觉信息间的不一致,而非单一模态。现有方法多依赖特征融合或跨模态注意力,难以有效捕捉不同表示层级的层次化语义矛盾。为此,本文提出HCIG(分层跨模态不一致图网络),利用图注意力网络在词、短语、全局三个层级建模跨模态不一致,并通过学习的层次注意力机制自适应融合表示。作为补充架构,提出GCCN(基于图的跨模态矛盾网络),采用矛盾感知池化实现高效多模态交互学习。模型在MMSD讽刺检测基准和MultiBully网络欺凌数据集上评估,包含全面消融实验与跨任务迁移测试。结果表明,HCIG在MMSD上取得85.74%准确率和85.29%宏F1,GCCN在MultiBully上达到最高宏F1(68.66%),HCIG在准确率(69.62%)与欺凌类F1(74.90%)上领先。结果证明,分层多粒度不一致建模比传统融合策略更有效,为社交媒体中的讽刺与欺凌检测提供稳健框架。
原文摘要 · Abstract (English)
Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multimodal approaches primarily rely on feature fusion or cross-modal attention, which may not effectively capture hierarchical semantic inconsistencies across different levels of representation. To address this limitation, this paper proposes HCIG (Hierarchical Cross-modal Incongruity Graph Network), a novel framework that models cross-modal incongruity at token, phrase, and global levels using graph attention networks and adaptively integrates these representations through a learned hierarchical attention mechanism. As a complementary architecture, we also introduce GCCN (Graph-based Cross-modal Contradiction Network), which performs graph-based reasoning using contradiction-aware pooling for efficient multimodal interaction learning. The proposed models are evaluated on the MMSD sarcasm benchmark and the MultiBully cyberbullying dataset, together with comprehensive ablation studies and cross-task transfer experiments. Experimental results demonstrate that HCIG achieves the best performance on MMSD with 85.74% accuracy and 85.29% macro-F1, while GCCN attains the highest macro-F1 (68.66%) on MultiBully and HCIG achieves the highest accuracy (69.62%) and bullying-class F1 (74.90%). The findings demonstrate that hierarchical multi-granularity incongruity modeling provides more effective multimodal reasoning than conventional fusion strategies, offering a robust framework for sarcasm and cyberbullying detection in social media.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。