arXiv:2607.19061cs.CVcs.AI2026-07

通过自适应提取隐藏视角,精准识别图文欺凌陷阱。

Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions

论文配图:Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions
图 1 · 摘自论文原文
  • 将隐性仇恨幻觉检测转为视觉检索问题,动态选择可信视角。
  • 在测试集上达93.2%准确率,显著超越原有方法与固定滤镜。
  • 适合安全审查、内容审核系统开发者使用,尤其应对隐蔽仇恨内容。

仇恨视觉错觉暴露了当前多模态安全系统的严重缺陷。在原始视角下,已有六种审核模型的准确率仅为20.9%至24.5%,九种顶尖视觉语言模型(VLMs)在幻觉感知提示下准确率仍不高于10.2%,多数隐藏仇恨未被发现。本文将隐藏仇恨幻觉检测建模为感知检索问题,提出自适应视角检索(Adaptive View Retrieval)框架。该框架构建图像与隐藏信息模板的互补视角库,自适应选择可信视角,检索隐藏信息身份,并校准恢复证据是否具有危害性。在冻结CLIP编码器条件下,该方法在HatefulIllusion数据集的保留测试集上达到93.2%的平衡准确率。其性能显著优于原始视角基线与固定单变换滤镜,在仇恨俚语、仇恨符号及可见度水平上均表现更优。同一设计还超越官方微调的CLIP基线,匹配或超过人类在IllusionMNIST、IllusionFashionMNIST和IllusionAnimals上的表现,并在SemVink协议下优于缩放预处理方案在HC-Bench的表现。结果表明,稳健的多模态审核必须先还原隐藏含义,再判断其危害性。

原文摘要 · Abstract (English)

Hateful optical illusions expose a serious gap in current multimodal safety systems. On original-view hateful illusions, previous work shows that six moderation classifiers achieve at most 20.9 to 24.5% accuracy and nine state-of-the-art VLMs remain at or below 10.2% with illusion-aware prompting, leaving most hidden hate undetected. We formulate hidden hateful illusion detection as a perceptual retrieval problem and propose Adaptive View Retrieval. This retrieve-and-calibrate framework assembles a complementary view bank for the image and hidden-message templates, adaptively selects which views to trust, retrieves hidden-message identities, and calibrates whether the recovered evidence is harmful. On HatefulIllusion with a frozen CLIP encoder, Adaptive View Retrieval reaches 93.2% balanced accuracy on the held-out test split. It substantially outperforms original-view baselines and fixed single-transform filters across hate slangs, hate symbols, and visibility levels. The same design also surpasses official fine-tuned CLIP baselines, matches or exceeds human performance on IllusionMNIST, IllusionFashionMNIST, and IllusionAnimals, and outperforms zoom-out preprocessing on HC-Bench under the SemVink protocol. Together, these results show that robust multimodal moderation requires recovering hidden meaning before deciding whether it is harmful.

多模态安全视觉幻觉内容审核自适应检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。