arXiv:2410.12700cs.CVcs.AI2024-10中稿 · ACM Multimedia 202…被引 4

轻量级方法让文生图模型生成更符合人类价值观的图像。

Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization

  • 通过优化轻量值编码器,动态注入价值原则控制生成内容。
  • 构建86k组图文偏好数据,实现高质量与价值观一致性平衡。
  • 无需大量参数更新,快速收敛,适合实际部署与伦理约束场景。

近期基于大规模数据训练的扩散模型已能生成媲美人类水平的图像,但常产生违背人类价值观的内容,如社会偏见和冒犯性信息。尽管大语言模型对齐研究进展显著,文本到图像生成模型的对齐问题仍鲜有探索。为此,我们提出轻量值优化(LiVO)方法,仅通过优化一个即插即用的值编码器,将指定价值准则融入输入提示,从而同时控制生成图像的语义与价值取向。我们设计了针对扩散模型的偏好优化损失,理论上逼近大语言模型中使用的Bradley-Terry模型,但提供更灵活的质量与价值权衡。为优化值编码器,我们自动构建了包含86,000个样本的文本-图像偏好数据集(提示、对齐图像、违反图像、价值原则)。该方法不更新多数模型参数,通过自适应从提示中选择价值,显著减少有害输出,收敛速度更快,超越多个强基线,迈出迈向伦理对齐文生图模型的重要一步。

原文摘要 · Abstract (English)

Recent advancements in diffusion models trained on large-scale data have enabled the generation of indistinguishable human-level images, yet they often produce harmful content misaligned with human values, e.g., social bias, and offensive content. Despite extensive research on Large Language Models (LLMs), the challenge of Text-to-Image (T2I) model alignment remains largely unexplored. Addressing this problem, we propose LiVO (Lightweight Value Optimization), a novel lightweight method for aligning T2I models with human values. LiVO only optimizes a plug-and-play value encoder to integrate a specified value principle with the input prompt, allowing the control of generated images over both semantics and values. Specifically, we design a diffusion model-tailored preference optimization loss, which theoretically approximates the Bradley-Terry model used in LLM alignment but provides a more flexible trade-off between image quality and value conformity. To optimize the value encoder, we also develop a framework to automatically construct a text-image preference dataset of 86k (prompt, aligned image, violating image, value principle) samples. Without updating most model parameters and through adaptive value selection from the input prompt, LiVO significantly reduces harmful outputs and achieves faster convergence, surpassing several strong baselines and taking an initial step towards ethically aligned T2I models.

文生图伦理对齐轻量优化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。