解决文本生成图像中的安全与质量矛盾,提出数据、方法、评估一体化方案。
The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation
- 构建双标注的大规模数据集LibraAlign-100K,缓解训练偏差
- 提出T2I-SPO算法,联合优化安全与质量偏好,提升生成能力
- 设计统一评分机制,细粒度衡量安全与生成性能的平衡
内容安全是文本到图像(T2I)模型的核心挑战,现有方法在安全与生成质量间存在严重权衡。我们指出,这一困境源于数据、方法与评估协议的系统性缺陷。为此,我们提出一体化安全对齐框架:首先,构建首个包含安全与质量双重标注的大型数据集LibraAlign-100K,克服训练信号偏倚;其次,提出协同偏好优化(T2I-SPO),扩展DPO范式,采用融合安全与质量的复合奖励函数,全面建模用户偏好;最后,引入统一对齐得分(Unified Alignment Score),以细粒度、非二值化方式公平评估安全与生成能力的平衡。大量实验表明,T2I-SPO在应对多种NSFW概念时达到领先的安全对齐水平,同时更有效地保持模型的生成质量与泛化能力。
原文摘要 · Abstract (English)
Content safety is a fundamental challenge for text-to-image (T2I) models, yet prevailing methods enforce a debilitating trade-off between safety and generation quality. We argue that mitigating this trade-off hinges on addressing systemic challenges in current T2I safety alignment across data, methods, and evaluation protocols. To this end, we introduce a unified framework for synergistic safety alignment. First, to overcome the flawed data paradigm that provides biased optimization signals, we develop LibraAlign-100K, the first large-scale dataset with dual annotations for safety and quality. Second, to address the myopic optimization of existing methods focus solely on safety reward, we propose Synergistic Preference Optimization (T2I-SPO), a novel alignment algorithm that extends the DPO paradigm with a composite reward function that integrates generation safety and quality to holistically model user preferences. Finally, to overcome the limitations of quality-agnostic and binary evaluation in current protocols, we introduce the Unified Alignment Score, a holistic, fine-grained metric that fairly quantifies the balance between safety and generative capability. Extensive experiments demonstrate that T2I-SPO achieves state-of-the-art safety alignment against a wide range of NSFW concepts, while better maintaining the model's generation quality and general capability
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。