PRJ框架让AI更懂图像隐性危害,像人一样分步判断内容安全。
PRJ: Perception-Retrieval-Judgement for Generated Images
- 分三步:看图描述、查相关知识、按规则判断,模拟人类认知过程。
- 在检测隐性有害内容上准确率超现有系统,支持细粒度分类。
- 适合内容审核、平台治理等需理解语境的AI安全场景。
生成式AI的快速发展带来了强大的创作能力,但也引发了内容安全的紧迫问题,如色情、暴力、仇恨符号、宣传信息及未经授权的艺术品模仿等。现有图像安全系统多依赖固定类别过滤,输出二值结果,难以理解上下文或识别对抗性诱导的隐蔽伤害。标准评估指标(如攻击成功率)也无法捕捉毒性语义严重性和动态演变。为此,我们提出感知-检索-判断(PRJ)框架,将毒性检测建模为结构化推理过程:首先将图像转化为语言描述(感知),然后检索与危害类别和特征相关的外部知识(检索),最后基于法律或规范准则进行毒性评估(判断)。该语言中心结构提升了对显性和隐性危害的检测能力,增强了可解释性与类别粒度。此外,我们设计了基于上下文毒性风险矩阵的动态评分机制,量化不同语义维度的有害程度。实验表明,PRJ在检测精度与鲁棒性上超越现有安全检查器,并首次实现分层级的毒性解释。
原文摘要 · Abstract (English)
The rapid progress of generative AI has enabled remarkable creative capabilities, yet it also raises urgent concerns regarding the safety of AI-generated visual content in real-world applications such as content moderation, platform governance, and digital media regulation. This includes unsafe material such as sexually explicit images, violent scenes, hate symbols, propaganda, and unauthorized imitations of copyrighted artworks. Existing image safety systems often rely on rigid category filters and produce binary outputs, lacking the capacity to interpret context or reason about nuanced, adversarially induced forms of harm. In addition, standard evaluation metrics (e.g., attack success rate) fail to capture the semantic severity and dynamic progression of toxicity. To address these limitations, we propose Perception-Retrieval-Judgement (PRJ), a cognitively inspired framework that models toxicity detection as a structured reasoning process. PRJ follows a three-stage design: it first transforms an image into descriptive language (perception), then retrieves external knowledge related to harm categories and traits (retrieval), and finally evaluates toxicity based on legal or normative rules (judgement). This language-centric structure enables the system to detect both explicit and implicit harms with improved interpretability and categorical granularity. In addition, we introduce a dynamic scoring mechanism based on a contextual toxicity risk matrix to quantify harmfulness across different semantic dimensions. Experiments show that PRJ surpasses existing safety checkers in detection accuracy and robustness while uniquely supporting structured category-level toxicity interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。