arXiv:2510.27196cs.CLcs.AI2025-10EMNLP

构建多视角评估框架,更公平地测试大模型对梗图危害性的理解能力。

MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models

  • 用多角色代理模拟不同语境,生成针对性分析任务。
  • 评估结果与人类偏好高度一致,显著降低评判偏差。
  • 适合研究AI安全、多模态理解及评测方法的学者使用。

社交媒体上梗图的泛滥要求多模态大语言模型(mLLMs)具备有效理解多模态危害性的能力。现有评估方法主要关注二分类检测准确率,难以反映危害性在不同语境下的深层解释差异。本文提出MemeArena,一种基于代理的竞技场式评估框架,可实现对mLLMs多模态危害性理解的上下文感知与无偏评估。具体而言,MemeArena通过模拟多样化的解释语境,生成能激发mLLMs观点特异性分析的评测任务。通过整合多方观点并达成评价共识,实现对mLLMs危害性理解能力的公平、无偏比较。大量实验表明,该框架有效降低了裁判代理的评估偏差,判断结果与人类偏好高度一致,为多模态危害性理解中可靠且全面的mLLM评估提供了重要洞见。代码与数据已公开于https://github.com/Lbotirx/MemeArena。

原文摘要 · Abstract (English)

The proliferation of memes on social media necessitates the capabilities of multimodal Large Language Models (mLLMs) to effectively understand multimodal harmfulness. Existing evaluation approaches predominantly focus on mLLMs' detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts. In this paper, we propose MemeArena, an agent-based arena-style evaluation framework that provides a context-aware and unbiased assessment for mLLMs' understanding of multimodal harmfulness. Specifically, MemeArena simulates diverse interpretive contexts to formulate evaluation tasks that elicit perspective-specific analyses from mLLMs. By integrating varied viewpoints and reaching consensus among evaluators, it enables fair and unbiased comparisons of mLLMs' abilities to interpret multimodal harmfulness. Extensive experiments demonstrate that our framework effectively reduces the evaluation biases of judge agents, with judgment results closely aligning with human preferences, offering valuable insights into reliable and comprehensive mLLM evaluations in multimodal harmfulness understanding. Our code and data are publicly available at https://github.com/Lbotirx/MemeArena.

多模态模型评估AI安全梗图理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。