arXiv:2605.28137cs.CVcs.LG2026-05

训练数据越不安全,生成图像就越危险,且无安全阈值。

No Safe Dose: How Training Data Drives Unsafe Image Generation

论文配图:No Safe Dose: How Training Data Drives Unsafe Image Generation
图 1 · 摘自论文原文
  • 用不同比例的不安全图像训练模型,仅改变数据构成。
  • 不安全图像占比从0%升至5%,输出不安全率从16.6%升至25.5%。
  • 即使数据干净,模型仍有16.6%不安全基线,适合关注安全性的研究者。

文本到图像模型在大规模数据上训练时不可避免会摄入不安全内容。尽管有人观察到输入输出的放大效应,但训练数据构成是否直接影响模型输出安全性仍不明确。我们通过隔离该变量:在仅含不安全图像比例不同(0%至9.6%)的数据集上(10万至800万样本),训练同一模型,并使用四个独立安全分类器评估生成结果。结果显示,输出不安全率从0%污染下的16.6%单调上升至5%污染下的25.5%。因子设计表明,影响因素是不安全图像的比例,而非绝对数量。零污染时仍存在16.6%不可消除的基线风险,表明冻结文本编码器等组件亦带来残留风险——通过文本编码器消融实验验证,使用SafeCLIP可将该基线降至9.6%,而剂量-反应关系在三种编码器中均持续存在。关键的是,安全过滤未导致FID、CLIPscore和ImageReward的性能下降。这些结果表明数据净化与文本编码器安全是互补且独立有效的干预手段。同时,剩余不安全水平引发对未来能力与组合性的研究思考。

原文摘要 · Abstract (English)

Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remains unclear whether and how training data composition directly drives model output safety or by other factors. We shed light on this question by isolating this variable: we train the same text-to-image model on datasets that differ \emph{only} in their fraction of unsafe images (0\% to 9.6\%), across several dataset scales (100K to 8M). Then we generate images with the resulting models, and evaluate them with four independent safety classifiers. Output unsafety rises monotonically from 16.6\% at 0\% contamination to 25.5\% at 5\%. A factorial design reveals that the \emph{proportion}, not the absolute count, of unsafe training images is the operative variable. The 16.6\% irreducible baseline at zero contamination implicates the other components, e.g. frozen text encoder, as a residual safety risk -- confirmed by a text encoder ablation showing that SafeCLIP reduces this floor to 9.6\%, while the dose-response effect persists across all three encoders tested. Critically, no quality degradation in terms of FID, CLIPscore and ImageReward accompanies safety filtering. These results establish that data curation and text encoder safety are complementary and independently effective interventions. At the same time, the remaining level of unsafety poses questions for future research about emerging capabilities and compositionality.

图像生成数据安全模型风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。