arXiv:2508.00445cs.CV2025-08中稿 · CVPR

自动检测并消除文本生成图像模型中的恶意偏见

AutoDebias: Automated Framework for Debiasing Text-to-Image Models

  • 用视觉语言模型识别触发式偏见模式
  • 91.6%准确率检测恶意模式,攻击成功率从90%降至几乎为零
  • 适合关注AI安全与公平性的研究者和开发者

文本生成图像(T2I)模型虽能生成高质量图像,却易受恶意后门攻击,注入有害偏见(如触发激活的性别或种族刻板印象)。现有去偏方法多针对自然统计偏见,难以应对此类故意且隐蔽的攻击。我们提出AutoDebias框架,无需预先了解攻击类型,即可自动识别并缓解这些恶意偏见。该框架利用视觉语言模型检测触发激活的视觉模式,并通过生成反向提示构建中和引导;再结合CLIP引导训练,打破有害关联的同时保持原模型图像质量与多样性。相较于针对自然偏见的方法,AutoDebias有效应对细微、被注入的刻板印象及多重交互攻击。我们在包含17种不同后门场景的新基准上进行评估,涵盖多个后门共存的复杂情况。结果表明,AutoDebias以91.6%的准确率检测恶意模式,将后门成功率从90%降至可忽略水平,同时保留原模型的视觉保真度。

原文摘要 · Abstract (English)

Text-to-Image (T2I) models generate high-quality images but are vulnerable to malicious backdoor attacks that inject harmful biases (e.g., trigger-activated gender or racial stereotypes). Existing debiasing methods, often designed for natural statistical biases, struggle with these deliberately and subtly injected attacks. We propose AutoDebias, a framework that automatically identifies and mitigates these malicious biases in T2I models without prior knowledge of the specific attack types. Specifically, AutoDebias leverages vision-language models to detect trigger-activated visual patterns and constructs neutralization guides by generating counter-prompts. These guides drive a CLIP-guided training process that breaks the harmful associations while preserving the original model's image quality and diversity. Unlike methods designed for natural bias, AutoDebias effectively addresses subtle, injected stereotypes and multiple interacting attacks. We evaluate the framework on a new benchmark covering 17 distinct backdoor scenarios, including challenging cases where multiple backdoors co-exist. AutoDebias detects malicious patterns with 91.6% accuracy and reduces the backdoor success rate from 90% to negligible levels, while preserving the visual fidelity of the original model.

文本生成图像偏见检测AI安全后门攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。