arXiv:2412.18216cs.CVcs.CL2024-12AAAI被引 16

让AI更懂规则,生成准确可解释的图像内容审核结果

ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation

  • 基于规则构建数据集,用多阶段提示增强标注
  • 分类准确率提升36.8%,解释质量提高26.6%
  • 适合需要合规、可解释审核的平台应用

网络上充斥着大量争议性内容,严重违反文化规范与儿童保护标准。传统图像内容审核模型难以满足多样化标准下的精准决策,而现有多模态大模型在通用规则化审核任务中常产生与人工审核不一致的分类与解释结果。为实现灵活、可解释且精准的图像内容审核,我们设计了一种新的规则数据集生成流程,将简明的人类定义规则分解,并通过精心设计的多阶段提示丰富短文本标注。所构建的ICM-Instruct数据集包含详细审核解释与问答对。基于此,我们提出了ICM-Assistant模型,在规则化审核框架下具备实际部署能力。该模型在多个来源上表现优异,平均分类准确率提升36.8%,解释质量提升26.6%,显著优于现有多模态大模型。代码与数据已开源。

原文摘要 · Abstract (English)

Controversial contents largely inundate the Internet, infringing various cultural norms and child protection standards. Traditional Image Content Moderation (ICM) models fall short in producing precise moderation decisions for diverse standards, while recent multimodal large language models (MLLMs), when adopted to general rule-based ICM, often produce classification and explanation results that are inconsistent with human moderators. Aiming at flexible, explainable, and accurate ICM, we design a novel rule-based dataset generation pipeline, decomposing concise human-defined rules and leveraging well-designed multi-stage prompts to enrich short explicit image annotations. Our ICM-Instruct dataset includes detailed moderation explanation and moderation Q-A pairs. Built upon it, we create our ICM-Assistant model in the framework of rule-based ICM, making it readily applicable in real practice. Our ICM-Assistant model demonstrates exceptional performance and flexibility. Specifically, it significantly outperforms existing approaches on various sources, improving both the moderation classification (36.8% on average) and moderation explanation quality (26.6% on average) consistently over existing MLLMs. Code/Data is available at https://github.com/zhaoyuzhi/ICM-Assistant.

图像审核多模态可解释性规则系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。