arXiv:2507.01702cs.CLcs.AI2025-07ACL被引 8

动态测试大模型对有害图文的推理能力,发现其真实短板。

AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness

  • 用多智能体协作动态生成挑战性梗图,持续探测模型弱点。
  • 实测多款大模型在识别有害内容上表现差异显著,暴露具体缺陷。
  • 适合研究安全评估、模型鲁棒性的研究人员使用。

社交媒体时代图文梗图泛滥,亟需大模型有效理解其有害性。现有评估基准依赖静态数据集和基于准确率的模型无关评测,难以跟上线上梗图的动态演变。为此,我们提出AdamMeme——一种灵活的代理式评估框架,通过多智能体协作,自适应地迭代更新梗图样本,持续探测大模型在识别有害性时的推理能力。实验表明,该框架能系统揭示不同目标模型的性能差异,提供细粒度的模型弱点分析。代码已开源:https://github.com/Lbotirx/AdamMeme。

原文摘要 · Abstract (English)

The proliferation of multimodal memes in the social media era demands that multimodal Large Language Models (mLLMs) effectively understand meme harmfulness. Existing benchmarks for assessing mLLMs on harmful meme understanding rely on accuracy-based, model-agnostic evaluations using static datasets. These benchmarks are limited in their ability to provide up-to-date and thorough assessments, as online memes evolve dynamically. To address this, we propose AdamMeme, a flexible, agent-based evaluation framework that adaptively probes the reasoning capabilities of mLLMs in deciphering meme harmfulness. Through multi-agent collaboration, AdamMeme provides comprehensive evaluations by iteratively updating the meme data with challenging samples, thereby exposing specific limitations in how mLLMs interpret harmfulness. Extensive experiments show that our framework systematically reveals the varying performance of different target mLLMs, offering in-depth, fine-grained analyses of model-specific weaknesses. Our code is available at https://github.com/Lbotirx/AdamMeme.

多模态模型评估安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。