arXiv:2511.19558cs.CRcs.AI2025-11被引 2

测试图像生成模型在微调后是否仍安全,发现常见方法易失效。

SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation

  • 提出SPQR多维评估框架,统一衡量安全、提示遵循性与鲁棒性。
  • 实测显示多数安全对齐方法在微调后会崩溃,暴露脆弱性。
  • 适合关注生成模型安全性的研究人员和产品开发者使用。

文本到图像扩散模型可能生成版权、不安全或隐私内容。安全对齐旨在抑制特定概念,但现有评估很少检验其在部署后常见良性微调(如LoRA个性化、风格/领域适配器)下的稳定性。我们研究当前安全方法在良性微调下的表现,发现频繁失效。真正的安全对齐必须能抵御此类后部署调整,因此提出SPQR基准(安全、提示遵循性、质量与鲁棒性)。SPQR是一个单分值指标,提供统一可复现的评估框架,通过单一排行榜分数衡量安全对齐模型在良性微调下保持安全性、可用性和鲁棒性的能力。我们进行多语言、领域特异及分布外分析,并进行类别级分解,识别安全对齐失败的场景,最终证明SPQR是文本到图像安全对齐技术的简洁而全面的基准。

原文摘要 · Abstract (English)

Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet evaluations seldom test whether safety persists under benign downstream fine-tuning routinely applied after deployment (e.g., LoRA personalization, style/domain adapters). We study the stability of current safety methods under benign fine-tuning and observe frequent breakdowns. As true safety alignment must withstand even benign post-deployment adaptations, we introduce the SPQR benchmark (Safety, Prompt adherence, Quality, and Robustness). SPQR is a single-scored metric that provides a unified, reproducible framework to evaluate how well safety-aligned diffusion models preserve safety, utility, and robustness under benign fine-tuning, by reporting a single leaderboard score to facilitate comparisons. We conduct multilingual, domain-specific, and out-of-distribution analyses, along with category-wise breakdowns, to identify when safety alignment fails after benign fine-tuning, ultimately showcasing SPQR as a concise yet comprehensive benchmark for T2I safety alignment techniques for T2I models.

安全对齐图像生成评估基准扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。