arXiv:2606.22958cs.LGcs.CV2026-06

提出无需训练的联合优化框架,提升文生图模型推理时对齐效果。

PG-MAP: Joint MAP Optimization for Inference-Time Alignment of Diffusion and Flow-Matching Models

论文配图:PG-MAP: Joint MAP Optimization for Inference-Time Alignment of Diffusion and Flow-Matching Models
图 1 · 摘自论文原文
  • 通过轨迹级吉布斯-最大后验优化,联合调整条件与隐变量
  • 在扩散模型上提升PickScore和美学评分,流匹配模型达91.9%准确率
  • 适用于不同生成模型,可配合引导策略实现更强性能

推理时对齐预训练文生图模型通常仅沿单一控制轴进行,如无分类器引导、注意力编辑或基于奖励的潜在扰动。这一限制阻碍了条件与潜在变量间联合依赖建模,并影响生成迁移能力。本文提出PG-MAP,一种无需训练的框架,将推理时对齐建模为条件$ c $与隐状态$ z_t $上的轨迹级吉布斯-MAP/近端能量优化,通过前向一致性耦合实现。该方法可选地由冻结的偏好奖励引导,支持扩散与流匹配模型的统一适配。在扩散模型(SD 1.5、SDXL)上,PG-MAP持续提升PickScore与美学评分,并可与调优后的无分类器引导结合,取得最佳综合表现。在流匹配模型(SD3.5-medium)上,框架退化为仅潜变量版本,实现91.9%的PickScore与75.7%的HPS胜率,受控实验排除了噪声干扰。人工评估确认其优于强基线,包括调优的CFG与计算量匹配的通用引导。此外,预言路由分析显示条件与潜在优化的重要性随提示类型变化,表明存在可通过每提示选择器进一步挖掘的潜力。

原文摘要 · Abstract (English)

Inference-time alignment of pretrained text-to-image models is typically performed along a single control axis, such as classifier-free guidance, attention editing, or reward-based latent perturbations. This limitation prevents modeling joint dependencies between conditioning and latent variables and hinders transfer across generative transports. We propose PG-MAP, a training-free framework that formulates inference-time alignment as a trajectory-level Gibbs-MAP / proximal energy optimization over the conditioning $c$ and latent state $z_t$ via a forward-consistency coupling, optionally guided by a frozen preference reward. This joint formulation enables coordinated updates across modalities while remaining compatible with both diffusion and flow-matching models through transport-specific adaptations. Across diffusion backbones (SD~1.5, SDXL), PG-MAP consistently improves alignment metrics such as PickScore and Aesthetic, and can be effectively combined with tuned classifier-free guidance to achieve the strongest overall performance. On flow-matching models (SD3.5-medium), the framework reduces to a latent-only variant, achieving $\mathbf{91.9\%}$ PickScore and $75.7\%$ HPS win rates against a static baseline, with controlled experiments ruling out noise-related artifacts. Human evaluations further confirm consistent preference over strong baselines, including tuned CFG and compute-matched universal guidance. Finally, an oracle-routing analysis shows that the relative importance of conditioning and latent optimization depends on prompt types, surfacing further headroom that a per-prompt selector could exploit.

生成模型对齐优化扩散模型流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。