提出无需提示词的模型审计方法,可高效检测扩散模型是否学会生成特定敏感概念。
What Lurks Within? Concept Auditing for Shared Diffusion Models at Scale
- 直接分析模型内部行为,不依赖生成图像或精心设计提示词。
- 在320个受控模型和771个真实模型上检测准确率超90%。
- 审计速度比现有方法快18到40倍,适合大规模模型共享场景。
扩散模型(DMs)已彻底改变文生图技术,使用户能基于文本提示生成高度真实且定制化的图像。随着参数高效微调(PEFT)技术的发展,用户可仅用少量资源对强大预训练模型进行定制。然而,这些微调后的模型在开放平台上的广泛共享引发了日益严重的伦理与法律担忧,因其可能无意或故意生成敏感或未经授权的内容。尽管监管关注不断上升,目前仍缺乏系统性工具用于部署前的模型审计。本文提出概念审计问题:判断一个微调后的扩散模型是否学会生成特定目标概念。现有方法多依赖提示词构造与输出图像分类,但存在提示不确定性、概念漂移和可扩展性差等关键缺陷。为此,我们提出无提示词、无图像审计(PAIA)框架,将扩散模型视为被检对象,直接分析其内部行为,避免使用优化提示或生成图像。我们在320个受控模型(基于精心筛选的概念数据集训练)和771个来自公共扩散模型共享平台的真实社区模型上评估了PAIA。结果表明,PAIA在检测准确率超过90%的同时,将审计时间缩短18至40倍。据我们所知,PAIA是首个可扩展且实用的扩散模型事前概念审计解决方案,为更安全、透明的扩散模型共享提供了实践基础。
原文摘要 · Abstract (English)
Diffusion models (DMs) have revolutionized text-to-image generation, enabling the creation of highly realistic and customized images from text prompts. With the rise of parameter-efficient fine-tuning (PEFT) techniques, users can now customize powerful pre-trained models using minimal computational resources. However, the widespread sharing of fine-tuned DMs on open platforms raises growing ethical and legal concerns, as these models may inadvertently or deliberately generate sensitive or unauthorized content. Despite increasing regulatory attention on generative AI, there are currently no practical tools for systematically auditing these models before deployment. In this paper, we address the problem of concept auditing: determining whether a fine-tuned DM has learned to generate a specific target concept. Existing approaches typically rely on prompt-based input crafting and output-based image classification but they suffer from critical limitations, including prompt uncertainty, concept drift, and poor scalability. To overcome these challenges, we introduce Prompt-Agnostic Image-Free Auditing (PAIA), a novel, model-centric concept auditing framework. By treating the DM as the object of inspection, PAIA enables direct analysis of internal model behavior, bypassing the need for optimized prompts or generated images. We evaluate PAIA on 320 controlled models trained with curated concept datasets and 771 real-world community models sourced from a public DM sharing platform. Evaluation results show that PAIA achieves over 90% detection accuracy while reducing auditing time by 18 - 40X compared to existing baselines. To our knowledge, PAIA is the first scalable and practical solution for pre-deployment concept auditing of diffusion models, providing a practical foundation for safer and more transparent diffusion model sharing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。