用细粒度属性树提升扩散模型绘画生成质量
Beyond Binary Preference: Aligning Diffusion Models to Fine-grained Criteria by Decoupling Attributes
- 构建树状细粒度属性体系,分解图像质量为正负属性
- 提出双阶段框架,先注入领域知识再优化多属性偏好
- 在绘画生成中显著提升与专家标准的一致性
扩散模型的后训练对齐通常依赖简化信号,如标量奖励或二元偏好,难以捕捉复杂的人类专业判断。为此,我们首先由领域专家构建分层细粒度评价标准,将图像质量分解为多个正负属性并组织成树结构。基于此,提出两阶段对齐框架:第一阶段通过监督微调将领域知识注入辅助扩散模型;第二阶段引入复杂偏好优化(CPO),扩展DPO以对齐目标扩散模型至非二元、分层标准。具体地,将对齐问题重构为同时最大化正属性概率、最小化负属性概率,并利用辅助扩散模型实现。我们在绘画生成领域应用该方法,基于自建标注数据集进行CPO训练。大量实验表明,CPO显著提升生成质量和与专家标准的一致性,为细粒度标准对齐开辟新路径。
原文摘要 · Abstract (English)
Post-training alignment of diffusion models relies on simplified signals, such as scalar rewards or binary preferences. This limits alignment with complex human expertise, which is hierarchical and fine-grained. To address this, we first construct a hierarchical, fine-grained evaluation criteria with domain experts, which decomposes image quality into multiple positive and negative attributes organized in a tree structure. Building on this, we propose a two-stage alignment framework. First, we inject domain knowledge to an auxiliary diffusion model via Supervised Fine-Tuning. Second, we introduce Complex Preference Optimization (CPO) that extends DPO to align the target diffusion to our non-binary, hierarchical criteria. Specifically, we reformulate the alignment problem to simultaneously maximize the probability of positive attributes while minimizing the probability of negative attributes with the auxiliary diffusion. We instantiate our approach in the domain of painting generation and conduct CPO training with an annotated dataset of painting with fine-grained attributes based on our criteria. Extensive experiments demonstrate that CPO significantly enhances generation quality and alignment with expertise, opening new avenues for fine-grained criteria alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。