无需训练即可提升图像生成与人类偏好对齐,通过动态调度多目标实现。
DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective Scheduling
- 利用注意力图增强早期扩散阶段语义对齐,无需额外训练。
- 引入动态多目标调度机制,适应不同生成阶段需求。
- 适用于各类预训练模型,兼容性强且效果稳定。
文本到图像的扩散模型对齐对于提升生成图像与人类偏好的一致性至关重要。尽管基于训练的方法受限于高计算成本和数据需求,但无需训练的对齐方法仍研究不足,且常因引导不准确而受限。本文提出一种即插即用的无需训练对齐方法 DyMO,用于在推理阶段对齐生成图像与人类偏好。除了基于文本的人类偏好评分外,还引入语义对齐目标,利用注意力图作为噪声图像中语义的有效反映,以增强早期扩散阶段的语义一致性。提出动态多目标调度与中间递归步骤,以适应不同生成阶段的需求。在多种预训练扩散模型与评估指标上的实验表明,该方法具有显著有效性与鲁棒性。
原文摘要 · Abstract (English)
Text-to-image diffusion model alignment is critical for improving the alignment between the generated images and human preferences. While training-based methods are constrained by high computational costs and dataset requirements, training-free alignment methods remain underexplored and are often limited by inaccurate guidance. We propose a plug-and-play training-free alignment method, DyMO, for aligning the generated images and human preferences during inference. Apart from text-aware human preference scores, we introduce a semantic alignment objective for enhancing the semantic alignment in the early stages of diffusion, relying on the fact that the attention maps are effective reflections of the semantics in noisy images. We propose dynamic scheduling of multiple objectives and intermediate recurrent steps to reflect the requirements at different steps. Experiments with diverse pre-trained diffusion models and metrics demonstrate the effectiveness and robustness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。