arXiv:2506.14903cs.CV2025-06被引 1

提升文生图模型对齐性,通过新方法增强安全与公平性

DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization

  • 引入DPO-Kernels,融合多种核函数和损失机制优化对齐
  • 在10万张图像对上验证,显著提升安全与非安全内容的分离能力
  • 适合关注生成模型安全、公平性的研究人员和开发者

对齐是确保文生图(T2I)模型忠实表达用户意图并维持安全与公平的关键。直接偏好优化(DPO)正从大语言模型扩展至T2I系统。本文提出DPO-Kernels,从三方面提升对齐:(i) 混合损失,结合嵌入式目标与传统概率损失以优化训练;(ii) 核化表示,采用径向基函数(RBF)、多项式与小波核实现更丰富的特征变换,增强安全/不安全输入的分离;(iii) 发散选择,拓展默认的KL正则化,引入Wasserstein与Rényi发散,提升稳定性与鲁棒性。我们构建了首个大规模基准DETONATE,包含约10万组精选图像对,分为优选与次优两类,涵盖种族、性别、残疾三类社会偏见维度。提示词来自仇恨言论数据集,图像由Stable Diffusion 3.5 Large、Stable Diffusion XL与Midjourney生成。此外,提出对齐质量指数(AQI),一种几何度量,量化潜在空间中安全与不安全图像激活的可分性,揭示隐藏漏洞。实证表明,DPO-Kernels通过重尾自正则化(HT-SR)保持强泛化界。所有数据与代码已公开。

原文摘要 · Abstract (English)

Alignment is crucial for text-to-image (T2I) models to ensure that generated images faithfully capture user intent while maintaining safety and fairness. Direct Preference Optimization (DPO), prominent in large language models (LLMs), is extending its influence to T2I systems. This paper introduces DPO-Kernels for T2I models, a novel extension enhancing alignment across three dimensions: (i) Hybrid Loss, integrating embedding-based objectives with traditional probability-based loss for improved optimization; (ii) Kernelized Representations, employing Radial Basis Function (RBF), Polynomial, and Wavelet kernels for richer feature transformations and better separation between safe and unsafe inputs; and (iii) Divergence Selection, expanding beyond DPO's default Kullback-Leibler (KL) regularizer by incorporating Wasserstein and R'enyi divergences for enhanced stability and robustness. We introduce DETONATE, the first large-scale benchmark of its kind, comprising approximately 100K curated image pairs categorized as chosen and rejected. DETONATE encapsulates three axes of social bias and discrimination: Race, Gender, and Disability. Prompts are sourced from hate speech datasets, with images generated by leading T2I models including Stable Diffusion 3.5 Large, Stable Diffusion XL, and Midjourney. Additionally, we propose the Alignment Quality Index (AQI), a novel geometric measure quantifying latent-space separability of safe/unsafe image activations, revealing hidden vulnerabilities. Empirically, we demonstrate that DPO-Kernels maintain strong generalization bounds via Heavy-Tailed Self-Regularization (HT-SR). DETONATE and complete code are publicly released.

文生图对齐优化安全生成偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。