通过内容令牌控制路由,提升扩散模型生成图像的偏好优化效果。
ToPO: Token-Conditioned Preference Routing for Attention-Based Latent Diffusion Models

- 基于分枝残差构建时空独立路由路径,实现细粒度偏好调节。
- 在五个SD-1.5指标上均优于Diffusion-DPO,SDXL也表现更优。
- 无需局部标签或奖励模型,适合追求生成质量的开发者使用。
成对偏好标签对完整图像进行排序,但扩散模型中的DPO方法将其影响扩展至多个空间与去噪时间坐标。针对基于注意力的噪声预测潜在扩散模型,本文提出ToPO(Token-Oriented Preference Optimization):在冻结参考去噪器中,基于分支平方残差构建每批次独立、可分离的时空路径。通过内容令牌调控空间因子的跨分支注意力,并引入无局部标签的像素中点排序辅助项。在共享更新调度的三种子重新训练中,ToPO在所有五个报告的SD-1.5指标上均优于Diffusion-DPO,且在HPSv2、ImageReward和CLIP评估下对SDXL同样表现更佳。在盲测的SDXL A/B实验中也获得更高原始胜率。结果限定于等更新量的U-Net协议,非等计算量对比。
原文摘要 · Abstract (English)
Pairwise preference labels rank complete images, yet Diffusion-DPO applies their effect over many spatial and denoising-time coordinates. For attention-based, noise-prediction latent diffusion, ToPO (Token-Oriented Preference Optimization) constructs a per-minibatch, detached, separable spatial-temporal route from branchwise squared-residual contrast in a frozen reference denoiser. Preferred-branch cross-attention uses content tokens to modulate the spatial factor, and an auxiliary pixel-midpoint ordering term is added without local labels or a learned reward model. In matched three-seed retrainings with a shared update schedule, ToPO has higher endpoint estimates than Diffusion-DPO on all five reported SD-1.5 metrics and on HPSv2, ImageReward, and CLIP for SDXL. It also receives larger raw win shares in an aggregate blind SDXL A/B study. These findings are scoped to the reported equal-update U-Net protocols rather than an equal-compute comparison.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。