arXiv:2504.14290cs.CV2025-04

解决文本生成图像中的安全与质量矛盾,提出数据、方法、评估一体化方案。

The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation

  • 构建双标注的大规模数据集LibraAlign-100K,缓解训练偏差
  • 提出T2I-SPO算法,联合优化安全与质量偏好,提升生成能力
  • 设计统一评分机制,细粒度衡量安全与生成性能的平衡

内容安全是文本到图像(T2I)模型的核心挑战,现有方法在安全与生成质量间存在严重权衡。我们指出,这一困境源于数据、方法与评估协议的系统性缺陷。为此,我们提出一体化安全对齐框架:首先,构建首个包含安全与质量双重标注的大型数据集LibraAlign-100K,克服训练信号偏倚;其次,提出协同偏好优化(T2I-SPO),扩展DPO范式,采用融合安全与质量的复合奖励函数,全面建模用户偏好;最后,引入统一对齐得分(Unified Alignment Score),以细粒度、非二值化方式公平评估安全与生成能力的平衡。大量实验表明,T2I-SPO在应对多种NSFW概念时达到领先的安全对齐水平,同时更有效地保持模型的生成质量与泛化能力。

原文摘要 · Abstract (English)

Content safety is a fundamental challenge for text-to-image (T2I) models, yet prevailing methods enforce a debilitating trade-off between safety and generation quality. We argue that mitigating this trade-off hinges on addressing systemic challenges in current T2I safety alignment across data, methods, and evaluation protocols. To this end, we introduce a unified framework for synergistic safety alignment. First, to overcome the flawed data paradigm that provides biased optimization signals, we develop LibraAlign-100K, the first large-scale dataset with dual annotations for safety and quality. Second, to address the myopic optimization of existing methods focus solely on safety reward, we propose Synergistic Preference Optimization (T2I-SPO), a novel alignment algorithm that extends the DPO paradigm with a composite reward function that integrates generation safety and quality to holistically model user preferences. Finally, to overcome the limitations of quality-agnostic and binary evaluation in current protocols, we introduce the Unified Alignment Score, a holistic, fine-grained metric that fairly quantifies the balance between safety and generative capability. Extensive experiments demonstrate that T2I-SPO achieves state-of-the-art safety alignment against a wide range of NSFW concepts, while better maintaining the model's generation quality and general capability

文本生成图像安全对齐偏好优化评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。