提出新基准与方法,让视觉语言模型真正理解仇恨迷因的语用机制。
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

- 构建25种修辞功能+10个目标群体的5000张仇恨迷因数据集
- 主流模型在新基准上准确率暴跌至接近随机水平
- 仅需500样本训练可显著提升模型性能,适合安全敏感场景
仇恨迷因检测对视觉语言模型仍具挑战性,因现有基准结构上为观察性,将修辞仇恨机制与目标群体特征混淆,阻碍对模型脆弱性的因果评估。为此,我们提出FBHM,一个基于功能的仇恨迷因系统化基准,沿两个正交维度构建:25种不同修辞功能与10个目标群体(共5000张迷因)。对前沿视觉语言模型的基准测试显示严重泛化差距:模型在标准数据集上表现优异,但在FBHM上性能急剧下降至接近随机水平,证明其依赖数据集特定启发式而非鲁棒的多模态推理。为高效弥补此差距,我们提出LSV(可学习调制向量),一种极低数据范式策略,在仅500个调制样本(50张基础迷因)上施加因果干预目标,使FBHM性能提升约30个宏平均F1点,优于上下文学习与参数高效微调,且不损害源域表现。
原文摘要 · Abstract (English)
Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding rhetorical hate mechanisms with target community features and preventing causal evaluation of model vulnerabilities. To address this, we introduce FBHM, a systematically curated benchmark of Functionality Based Hateful Memes constructed along two orthogonal axes: 25 distinct rhetorical functionalities and 10 target communities (5,000 memes total). Benchmarking state-of-the-art VLMs reveals a severe generalization gap: models highly accurate on standard datasets catastrophically drop to near-random performance on FBHM, proving they exploit dataset-specific heuristics rather than robust multimodal reasoning. To efficiently close this gap, we propose LSV (learnable steering vectors), an ultra-low data regime strategy that applies a causal intervention objective on as few as 500 steering samples (50 unique base memes), boosting FBHM performance by ~30 Macro-F1 points while outperforming in-context learning and PEFT without degrading source-domain performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。