arXiv:2602.01101cs.CV2026-02中稿 · WWW2026被引 3

提出共享表示学习方法,让图像识别更抗缺失文本

Robust Harmful Meme Detection under Missing Modalities via Shared Representation Learning

  • 通过独立投影各模态,学习跨模态共享表示
  • 在文本缺失时仍保持检测性能,优于现有方法
  • 适合实际中图文不全的有害梗图检测场景

网络梗图是传播政治、心理与社会文化思想的强大工具,但也可能被用于针对特定个体或群体的仇恨宣传。尽管已有研究致力于设计新型检测方法,但这些方法通常依赖于图文完整的数据。在真实场景中,由于光学字符识别(OCR)质量差等原因,文本模态常缺失,导致现有方法对缺失信息敏感,性能下降。为此,本文首次系统研究了有害梗图检测在模态不完整数据下的表现。我们提出一种新基线方法,通过独立投影各模态以学习共享表示,从而在模态缺失时仍能有效利用视觉特征。在两个基准数据集上的实验表明,当文本缺失时,该方法显著优于现有方法,且增强了视觉特征的融合能力,降低了对文本的依赖,提升了鲁棒性。本工作为有害梗图检测在真实世界的应用,特别是在模态缺失场景下,迈出了关键一步。

原文摘要 · Abstract (English)

Internet memes are powerful tools for communication, capable of spreading political, psychological, and sociocultural ideas. However, they can be harmful and can be used to disseminate hate toward targeted individuals or groups. Although previous studies have focused on designing new detection methods, these often rely on modal-complete data, such as text and images. In real-world settings, however, modalities like text may be missing due to issues like poor OCR quality, making existing methods sensitive to missing information and leading to performance deterioration. To address this gap, in this paper, we present the first-of-its-kind work to comprehensively investigate the behavior of harmful meme detection methods in the presence of modal-incomplete data. Specifically, we propose a new baseline method that learns a shared representation for multiple modalities by projecting them independently. These shared representations can then be leveraged when data is modal-incomplete. Experimental results on two benchmark datasets demonstrate that our method outperforms existing approaches when text is missing. Moreover, these results suggest that our method allows for better integration of visual features, reducing dependence on text and improving robustness in scenarios where textual information is missing. Our work represents a significant step forward in enabling the real-world application of harmful meme detection, particularly in situations where a modality is absent.

多模态图像识别抗缺失安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。