不生成图像即可快速检测恶意概念,防止有害内容传播。
Detecting Malicious Concepts without Image Generation in AI-Generated Content (AIGC)
- 仅分析概念文件,无需生成图像进行检测。
- 支持精准匹配与模糊检测两种模式,准确识别恶意概念。
- 适合平台方快速筛查上传概念,提升安全效率。
文本到图像生成任务在实践中已取得巨大成功,新兴的概念生成模型能够创建高度个性化和定制化的内容。用户对概念生成的热情迅速上升,概念共享平台也不断涌现。概念所有者可能上传恶意概念,并用非恶意的文本描述和示例图像伪装,诱导用户下载并生成有害内容。平台亟需一种快速判断概念是否恶意的方法,以阻止恶意概念传播。然而,单纯依赖生成概念图像来判断恶意性,不仅耗时且消耗大量计算资源。随着平台上上传和下载的概念数量持续增加,该方法变得不切实际,并存在生成恶意内容的风险。本文提出 Concept QuickLook,首个将恶意概念检测系统化纳入研究的工作,仅基于概念文件完成检测,无需生成任何图像。我们定义了恶意概念,并设计了两种检测模式:概念匹配与模糊检测。大量实验表明,Concept QuickLook 能有效检测恶意概念,在概念共享平台中具备实用性。我们还设计了鲁棒性实验进一步验证方案的有效性。期望本工作能开启恶意概念检测的研究,并提供启发。
原文摘要 · Abstract (English)
The task of text-to-image generation has achieved tremendous success in practice, with emerging concept generation models capable of producing highly personalized and customized content. Fervor for concept generation is increasing rapidly among users, and platforms for concept sharing have sprung up. The concept owners may upload malicious concepts and disguise them with non-malicious text descriptions and example images to deceive users into downloading and generating malicious content. The platform needs a quick method to determine whether a concept is malicious to prevent the spread of malicious concepts. However, simply relying on concept image generation to judge whether a concept is malicious requires time and computational resources. Especially, as the number of concepts uploaded and downloaded on the platform continues to increase, this approach becomes impractical and poses a risk of generating malicious content. In this paper, we propose Concept QuickLook, the first systematic work to incorporate malicious concept detection into research, which performs detection based solely on concept files without generating any images. We define malicious concepts and design two operational modes for detection: concept matching and fuzzy detection. Extensive experiments demonstrate that the proposed Concept QuickLook can detect malicious concepts and demonstrate practicality in concept sharing platforms. We also design robustness experiments to further validate the effectiveness of the solution. We hope this work can initiate malicious concept detection tasks and provide some inspiration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。