arXiv:2503.12560cs.CL2025-03AAAI被引 19

通过多粒度融合视觉与文本线索,提升表情包语义理解能力

Multi-Granular Multimodal Clue Fusion for Meme Understanding

  • 分层级提取图像物体特征,捕捉细粒度隐喻线索
  • 设计跨模态交互机制,强化图文弱关联的表示学习
  • 双语数据集上各项任务准确率提升超3.5%,适合多模态理解研究者

随着社交媒体平台的持续涌现,多模态表情包理解(MMU)任务日益受到关注。该任务旨在通过隐喻识别、情感分析、意图检测和攻击性检测等任务,从多角度理解表情包含义。尽管已有进展,但仍受限于细粒度视觉隐喻线索丢失以及图文弱相关性被忽视。为此,本文提出多粒度多模态线索融合模型(MGMCF),以推进MMU。首先,设计基于对象级别的语义挖掘模块,提取图像中的对象级特征线索,实现细粒度特征提取,增强对隐喻细节与语义的理解能力。其次,提出全新的全局-局部跨模态交互模型,解决文本与图像间的弱相关性问题,通过双向跨模态注意力机制促进全局多模态上下文线索与局部单模态特征线索的有效交互,增强表示。最后,设计双语义引导训练策略,提升模型在语义空间中对多模态表示的理解与对齐能力。在广泛使用的双语MET-MEME数据集上的实验表明,相比当前最优基线,该方法显著提升性能:攻击性检测任务精确率提升8.14%,隐喻识别、情感分析和意图检测任务的准确率分别提升3.53%、3.89%和3.52%。深入分析进一步验证了该方法的有效性与潜力。

原文摘要 · Abstract (English)

With the continuous emergence of various social media platforms frequently used in daily life, the multimodal meme understanding (MMU) task has been garnering increasing attention. MMU aims to explore and comprehend the meanings of memes from various perspectives by performing tasks such as metaphor recognition, sentiment analysis, intention detection, and offensiveness detection. Despite making progress, limitations persist due to the loss of fine-grained metaphorical visual clue and the neglect of multimodal text-image weak correlation. To overcome these limitations, we propose a multi-granular multimodal clue fusion model (MGMCF) to advance MMU. Firstly, we design an object-level semantic mining module to extract object-level image feature clues, achieving fine-grained feature clue extraction and enhancing the model's ability to capture metaphorical details and semantics. Secondly, we propose a brand-new global-local cross-modal interaction model to address the weak correlation between text and images. This model facilitates effective interaction between global multimodal contextual clues and local unimodal feature clues, strengthening their representations through a bidirectional cross-modal attention mechanism. Finally, we devise a dual-semantic guided training strategy to enhance the model's understanding and alignment of multimodal representations in the semantic space. Experiments conducted on the widely-used MET-MEME bilingual dataset demonstrate significant improvements over state-of-the-art baselines. Specifically, there is an 8.14% increase in precision for offensiveness detection task, and respective accuracy enhancements of 3.53%, 3.89%, and 3.52% for metaphor recognition, sentiment analysis, and intention detection tasks. These results, underpinned by in-depth analyses, underscore the effectiveness and potential of our approach for advancing MMU.

多模态理解表情包分析跨模态融合隐喻识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。