arXiv:2502.03512cs.AI2025-02

提出新基准,评估文本生成图像中的矛盾目标对齐问题。

YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment

  • 构建六组对立设计目标的系统性评估框架。
  • 包含人类提示、对齐与非对齐图像及矛盾解释数据集。
  • 为提升图像生成忠实度与美学伦理提供优化方向。

文本到图像(T2I)系统中的精确对齐至关重要,以确保生成图像不仅准确反映用户意图,还符合严格的伦理与审美标准。谷歌Gemini事件中因输出错位引发公众强烈反弹,凸显了鲁棒对齐机制的迫切需求。相比之下,大语言模型(LLM)已在对齐方面取得显著进展。基于此,研究者希望将类似直接偏好优化(DPO)的技术应用于T2I系统,以提升图像生成的保真度与可靠性。本文提出YinYangAlign,一个先进的基准测试框架,系统量化T2I系统的对齐保真度,涵盖六个基本且内在矛盾的设计目标。每组目标代表图像生成中的根本张力,如在遵循用户提示与创造性修改之间取得平衡,或在多样性与视觉一致性之间权衡。YinYangAlign包含详细的公理化数据集,包含人类提示、对齐(选择)结果、非对齐(拒绝)的AI生成输出,以及底层矛盾的解释。

原文摘要 · Abstract (English)

Precise alignment in Text-to-Image (T2I) systems is crucial to ensure that generated visuals not only accurately encapsulate user intents but also conform to stringent ethical and aesthetic benchmarks. Incidents like the Google Gemini fiasco, where misaligned outputs triggered significant public backlash, underscore the critical need for robust alignment mechanisms. In contrast, Large Language Models (LLMs) have achieved notable success in alignment. Building on these advancements, researchers are eager to apply similar alignment techniques, such as Direct Preference Optimization (DPO), to T2I systems to enhance image generation fidelity and reliability. We present YinYangAlign, an advanced benchmarking framework that systematically quantifies the alignment fidelity of T2I systems, addressing six fundamental and inherently contradictory design objectives. Each pair represents fundamental tensions in image generation, such as balancing adherence to user prompts with creative modifications or maintaining diversity alongside visual coherence. YinYangAlign includes detailed axiom datasets featuring human prompts, aligned (chosen) responses, misaligned (rejected) AI-generated outputs, and explanations of the underlying contradictions.

文本生成图像对齐评估多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。