arXiv:2602.03822cs.CL2026-02中稿 · the The Web Confer…被引 2

破解网络恶搞的隐喻陷阱,让AI看懂梗背后的伤害

They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References

  • 用知识图谱增强多模态理解,补足文化背景盲区
  • 通过轻量适配器提升判断精准度,减少误判讽刺与攻击
  • 自动生成逐层解释,让AI决策过程透明可查

基于表情包的社交欺凌检测面临挑战,因有害意图常依赖隐含的文化符号与跨模态不一致。现有方法(如融合模型、大视觉语言模型上下文学习)受限于三点:一、文化盲区(缺失象征语境),二、边界模糊(讽刺与攻击混淆),三、可解释性差(模型推理不透明)。我们提出CROSS-ALIGN+,一个三阶段框架系统解决上述问题:(1) 第一阶段通过整合ConceptNet、Wikidata和Hatebase的结构化知识,丰富多模态表示;(2) 第二阶段采用参数高效的LoRA适配器,精炼决策边界;(3) 第三阶段生成级联式解释,增强可解释性。在五个基准数据集和八种大视觉语言模型上实验证明,CROSS-ALIGN+持续优于现有最优方法,相对F1最高提升17%,并为每项判断提供可解释依据。

原文摘要 · Abstract (English)

Meme-based social abuse detection is challenging because harmful intent often relies on implicit cultural symbolism and subtle cross-modal incongruence. Prior approaches, from fusion-based methods to in-context learning with Large Vision-Language Models (LVLMs), have made progress but remain limited by three factors: i) cultural blindness (missing symbolic context), ii) boundary ambiguity (satire vs. abuse confusion), and iii) lack of interpretability (opaque model reasoning). We introduce CROSS-ALIGN+, a three-stage framework that systematically addresses these limitations: (1) Stage I mitigates cultural blindness by enriching multimodal representations with structured knowledge from ConceptNet, Wikidata, and Hatebase; (2) Stage II reduces boundary ambiguity through parameter-efficient LoRA adapters that sharpen decision boundaries; and (3) Stage III enhances interpretability by generating cascaded explanations. Extensive experiments on five benchmarks and eight LVLMs demonstrate that CROSS-ALIGN+ consistently outperforms state-of-the-art methods, achieving up to 17% relative F1 improvement while providing interpretable justifications for each decision.

内容安全多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。