arXiv:2502.12579cs.CV2025-02ICML被引 5

通过融合人类偏好与测试时采样,提升文生图模型对齐性与质量。

CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation

  • 分离建模偏好与非偏好分布,利用代理提示采样挖掘双重信息。
  • 仅用少量高质量微调数据即达顶尖性能,数据效率显著。
  • 适合追求高对齐性与生成质量的文生图研究与应用开发者。

扩散模型已成为文本到图像生成的主流方法。人类偏好对齐和无分类器引导等关键组件对保障生成质量至关重要。然而,当前模型独立应用这些技术仍面临文本-图像对齐弱、生成质量不高、与人类审美标准不一致等挑战。本文首次探索将人类偏好对齐与测试时采样协同优化,提出CHATS(Combining Human-Aligned optimization and Test-time Sampling)框架。该框架分别建模偏好与非偏好分布,并采用基于代理提示的采样策略,有效利用两类分布中的有用信息。实验表明,CHATS在极小规模高质量微调数据下即可实现卓越性能,超越传统偏好对齐方法,在多个标准基准上达到新最优水平。

原文摘要 · Abstract (English)

Diffusion models have emerged as a dominant approach for text-to-image generation. Key components such as the human preference alignment and classifier-free guidance play a crucial role in ensuring generation quality. However, their independent application in current text-to-image models continues to face significant challenges in achieving strong text-image alignment, high generation quality, and consistency with human aesthetic standards. In this work, we for the first time, explore facilitating the collaboration of human performance alignment and test-time sampling to unlock the potential of text-to-image models. Consequently, we introduce CHATS (Combining Human-Aligned optimization and Test-time Sampling), a novel generative framework that separately models the preferred and dispreferred distributions and employs a proxy-prompt-based sampling strategy to utilize the useful information contained in both distributions. We observe that CHATS exhibits exceptional data efficiency, achieving strong performance with only a small, high-quality funetuning dataset. Extensive experiments demonstrate that CHATS surpasses traditional preference alignment methods, setting new state-of-the-art across various standard benchmarks.

文生图扩散模型人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。