arXiv:2605.02940cs.LGcs.AI2026-05

零样本多智能体框架,让模型像侦探一样解析恶搞图的潜在危害。

PrismAgent: Illuminating Harm in Memes via a Zero-Shot Interpretable Multi-Agent Framework

论文配图:PrismAgent: Illuminating Harm in Memes via a Zero-Shot Interpretable Multi-Agent Framework
图 1 · 摘自论文原文
  • 将识别恶搞图危害类比为破案,四智能体分阶段协作分析。
  • 在三个数据集上超越现有零样本方法,无需标注数据即可精准检测。
  • 每步推理可解释,适合需要透明决策的可信AI应用。

恶搞图的快速传播使有害内容检测愈发重要,有效识别可遏制虚假信息扩散。然而,现有方法严重依赖大量标注数据,导致训练成本高且泛化能力弱。为此,我们提出PrismAgent,一种零样本、多智能体、可解释的框架。该框架将任务类比为刑事案件调查,采用四个专用智能体分别负责分析、调查、起诉和判决四个阶段。分析阶段,分析师智能体在善意与恶意假设下改写每张恶搞图,探测其潜在意图;调查阶段,调查员智能体从无标注数据集中检索证据,构建恶搞图及其变体的上下文解释;起诉阶段,检察官智能体将原图分别与三种解释配对,进行独立初步判断;判决阶段,裁判员智能体综合所有证据作出最终裁决。此外,PrismAgent的显式多阶段推理链使其天然可解释,每个中间步骤均有明确说明,而非仅输出最终结果。在三个公开数据集上的大量实验表明,PrismAgent显著优于现有零样本检测方法。

原文摘要 · Abstract (English)

The rapid spread of memes makes harmful content detection increasingly crucial, as effective identification can curb the circulation of misinformation. However, existing methods rely heavily on high-volume annotated data, which leads to substantial training costs and limited generalization. To address these challenges, we propose PrismAgent, a zero-shot, multi-agent, interpretable framework. PrismAgent conceptualizes this task as a criminal case investigation, employing four specialized agents responsible for the analysis, investigation, prosecution, and judgment stages within a structured collaborative workflow. In the first stage, the analyst agent paraphrases each meme under benevolent and malicious assumptions to probe its underlying intent. The investigator agent then retrieves supporting evidence from an unannotated dataset and constructs contextual interpretations for the meme and its variants. Next, the prosecutor agent performs three independent preliminary judgments by pairing the original meme with each of the three interpretations. Finally, the judge agent deliberates across all evidence to render a final verdict. Moreover, PrismAgent's explicit multi-stage reasoning chain makes the model inherently interpretable, as every intermediate step is explicitly explained rather than only producing a final detection result. Extensive experiments on three public datasets show that PrismAgent significantly outperforms existing zero-shot detection methods.

多智能体可解释性零样本图像安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。