arXiv:2508.01151cs.CVcs.AI2025-08被引 1

让AI画画更懂用户偏好,自动调节安全尺度。

Personalized Safety Alignment for Text-to-Image Diffusion Models

  • 用用户画像动态调整生成内容的安全阈值
  • 在宽松设置下提升画质,在严格设置下更强抑制风险
  • 适合需要个性化内容安全的AI绘画应用

文本到图像扩散模型虽已革新视觉生成,但部署受限于统一僵化的安全机制,无法体现年龄、文化或个人信仰带来的多元偏好差异。为此,我们提出个性化安全对齐(PSA)框架,将生成安全从静态过滤转向用户条件化适配。构建Sage大规模数据集,涵盖1000个模拟用户画像,覆盖传统数据集常遗漏的复杂风险。通过参数高效交叉注意力适配器整合用户画像,PSA动态调节生成以匹配个体敏感度。大量实验表明,PSA实现安全与质量的精准平衡:在宽松配置下缓解过度保守约束,提升视觉保真度;在严格配置下实现领先的风险抑制效果,显著优于静态基线。此外,其指令遵循能力优于提示工程方法,确立个性化是构建自适应、以用户为中心且负责任生成式AI的关键方向。代码、数据与模型已公开于https://github.com/M-E-AGI-Lab/PSAlign。

原文摘要 · Abstract (English)

Text-to-image diffusion models have revolutionized visual content generation, yet their deployment is hindered by a fundamental limitation: safety mechanisms enforce rigid, uniform standards that fail to reflect diverse user preferences shaped by age, culture, or personal beliefs. To address this, we propose Personalized Safety Alignment (PSA), a framework that transitions generative safety from static filtration to user-conditioned adaptation. We introduce Sage, a large-scale dataset capturing diverse safety boundaries across 1,000 simulated user profiles, covering complex risks often missed by traditional datasets. By integrating these profiles via a parameter-efficient cross-attention adapter, PSA dynamically modulates generation to align with individual sensitivities. Extensive experiments demonstrate that PSA achieves a calibrated safety-quality trade-off: under permissive profiles, it relaxes over-cautious constraints to enhance visual fidelity, while under restrictive profiles, it enforces state-of-the-art suppression, significantly outperforming static baselines. Furthermore, PSA exhibits superior instruction adherence compared to prompt-engineering methods, establishing personalization as a vital direction for creating adaptive, user-centered, and responsible generative AI. Our code, data, and models are publicly available at https://github.com/M-E-AGI-Lab/PSAlign.

图像生成安全对齐个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。