动态融合图文信息,提升恶搞内容识别准确率
DARC-CLIP: Dynamic Adaptive Refinement with Cross-Attention for Meme Understanding

- 引入双向注意力与动态适配模块,实现图文交互优化
- 在PrideMM上仇恨检测F1提升6.84,AUROC提升4.18
- 适合社交媒体有害内容检测、多模态分析研究者
恶搞图片通过视觉与文本信号的互动传递意义,常包含幽默、反讽和冒犯性内容。识别其中的有害或敏感信息需精准建模多模态线索。现有基于CLIP的方法依赖静态融合,难以捕捉模态间的细粒度依赖。我们提出DARC-CLIP,一种基于CLIP的自适应多模态融合框架,采用分层精炼结构。DARC-CLIP引入自适应交叉注意力精炼器(ACAR)实现双向信息对齐,以及任务感知的动态特征适配器(DFA)进行信号调制。我们在PrideMM基准上评估,该数据集涵盖仇恨、目标、立场和幽默分类,并进一步在CrisisHateMM数据集测试泛化能力。DARC-CLIP在各项任务中表现优异,仇恨检测方面相较最强基线实现+4.18 AUROC和+6.84 F1的显著提升。消融实验确认ACAR与DFA是性能提升的核心因素。结果表明,自适应跨信号精炼是社会敏感分类中多模态分析的有效策略。
原文摘要 · Abstract (English)
Memes convey meaning through the interaction of visual and textual signals, often combining humor, irony, and offense in subtle ways. Detecting harmful or sensitive content in memes requires accurate modeling of these multimodal cues. Existing CLIP-based approaches rely on static fusion, which struggles to capture fine grained dependencies between modalities. We propose DARC-CLIP, a CLIP-based framework for adaptive multimodal fusion with a hierarchical refinement stack. DARC-CLIP introduces Adaptive Cross-Attention Refiners to for bidirectional information alignment and Dynamic Feature Adapters for task-sensitive signal adaptation. We evaluate DARC-CLIP on the PrideMM benchmark, which includes hate, target, stance, and humor classification, and further test generalization on the CrisisHateMM dataset. DARC-CLIP achieves highly competitive classification accuracy across tasks, with significant gains of +4.18 AUROC and +6.84 F1 in hate detection over the strongest baseline. Ablation studies confirm that ACAR and DFA are the main contributors to these gains. These results show that adaptive cross-signal refinement is an effective strategy for multimodal content analysis in socially sensitive classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。