arXiv:2505.19010cs.CVcs.CL2025-05被引 2

通过细粒度跨模态注意力与通道门控,提升多模态有害内容检测效果

Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection

  • 引入通道级门控与协同注意力,实现图文特征的动态融合
  • 在MIMIC和SemEval数据集上达到当前最优性能
  • 适合需要精准跨模态对齐的社交媒体内容审核场景

多模态学习已成为关键研究方向,融合文本与视觉信息可显著提升分类、检索和场景理解等任务性能。尽管大模型取得进展,现有方法仍存在跨模态交互不足和融合策略僵化的问题,难以充分发挥不同模态的互补优势。为此,本文提出Co-AttenDWG,即协同注意力与通道级门控及专家融合机制。首先将文本与视觉特征映射至共享嵌入空间,通过专用协同注意力机制实现模态间的细粒度同步交互;进一步引入通道级门控网络,在通道层面自适应调节特征贡献,突出关键信息。同时,双路径编码器独立优化模态特异性表示,并通过额外的交叉注意力层实现更深层对齐。最终特征经专家融合模块结合学习到的门控与自注意力机制,生成鲁棒统一表征。在MIMIC和SemEval Memotion 1.0数据集上的实验表明,Co-AttenDWG达到最先进性能,展现出优越的跨模态对齐能力,验证了其在多样化多模态应用中的有效性。

原文摘要 · Abstract (English)

Multi-modal learning has emerged as a crucial research direction, as integrating textual and visual information can substantially enhance performance in tasks such as classification, retrieval, and scene understanding. Despite advances with large pre-trained models, existing approaches often suffer from insufficient cross-modal interactions and rigid fusion strategies, failing to fully harness the complementary strengths of different modalities. To address these limitations, we propose Co-AttenDWG, co-attention with dimension-wise gating, and expert fusion. Our approach first projects textual and visual features into a shared embedding space, where a dedicated co-attention mechanism enables simultaneous, fine-grained interactions between modalities. This is further strengthened by a dimension-wise gating network, which adaptively modulates feature contributions at the channel level to emphasize salient information. In parallel, dual-path encoders independently refine modality-specific representations, while an additional cross-attention layer aligns the modalities further. The resulting features are aggregated via an expert fusion module that integrates learned gating and self-attention, yielding a robust unified representation. Experimental results on the MIMIC and SemEval Memotion 1.0 datasets show that Co-AttenDWG achieves state-of-the-art performance and superior cross-modal alignment, highlighting its effectiveness for diverse multi-modal applications.

多模态有害内容检测协同注意力门控机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。