用视觉-文本注意力融合,精准识别孟加拉语政治梗图意图
Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

- 通过跨模态注意力对齐图文语义,实现视觉区域与文本词元的精准匹配
- 在PoliMemeDecode1数据集上达到0.94的宏平均F1,超越单模态与拼接方法
- 适用于低资源语言的政治情绪分析,模型可解释性强
社交媒体上多模态内容的自动化分析已成为理解公众情绪与信息传播的关键任务。然而,由于视觉线索与嵌入式、常为风格化的文字之间存在复杂交互,特别是对孟加拉语等低资源语言而言,互联网梗图的分类仍具挑战性。本文针对孟加拉语梗图中的政治意图检测,提出多模态交叉注意力融合框架。首先利用视觉-语言模型从嘈杂的梗图图像中提取高保真OCR文本;随后编码视觉与文本特征,并通过跨模态多头注意力机制进行融合,实现语义词元与视觉区域的对齐。同时,我们探索了领域特定政治词汇表作为知识先验的集成效果。在PoliMemeDecode1数据集上的实验表明,该注意力融合方法显著优于单模态基线和标准拼接方法,达到约0.94的宏观F1,为当前最优水平。可解释性分析进一步证实模型能有效将文本语义锚定于视觉证据。
原文摘要 · Abstract (English)
Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age. However, classifying internet memes remains computationally challenging due to the intricate interplay between visual cues and embedded, often stylized, text, particularly in low-resource languages like Bengali Language. This paper addresses the detection of political intent in Bengali memes by introducing Multimodal Cross-Attention Fusion framework. We first leverage a Vision-Language Model to extract high-fidelity OCR text from noisy meme images. Subsequently, we encode visual and textual features and synthesize them through a cross-modal multi-head attention mechanism that aligns semantic tokens with visual regions. We also investigate the integration of a domain-specific political lexicon as a knowledge prior. Experimental evaluation on the PoliMemeDecode1 dataset shows that our attention-based fusion significantly outperforms unimodal baselines and standard concatenation methods, achieving a state-of-the-art Macro-F1 of approximately 0.94. Interpretability analyzes further confirm that the model effectively learns to ground textual semantics in visual evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。