arXiv:2601.15968cs.CV2026-01被引 3

用超网络实现扩散模型高效测试时对齐,提升生成质量与提示一致性。

HyperAlign: Hypernetwork for Efficient Test-Time Alignment of Diffusion Models

  • 通过超网络动态生成低秩适配权重,调节去噪轨迹以匹配目标奖励。
  • 在Stable Diffusion和FLUX上显著优于现有方法,兼顾语义一致性和视觉质量。
  • 适合需要快速适应新提示又不牺牲多样性的生成应用开发者。

扩散模型对齐旨在通过增强生成结果与文本提示的语义一致性及整体视觉质量,弥合生成内容与人类偏好的差距。现有对齐方法面临挑战:测试时方法虽能实现输入特异性适应,但计算开销大且易欠优化;微调方法则可能引发奖励过拟合并损失生成多样性。为此,我们提出HyperAlign,一种训练超网络以实现高效、有效的测试时对齐框架。不同于直接修改潜在状态,HyperAlign动态生成依赖输入与状态的低秩适配权重,调控去噪轨迹以逼近目标奖励。我们设计了多种不同粒度的HyperAlign变体,平衡对齐效果与计算效率。超网络通过带偏好数据正则化的奖励目标进行优化,缓解奖励劫持问题。我们在多个生成范式(包括Stable Diffusion和FLUX)上评估,结果表明其在语义一致性和视觉质量上均显著优于现有方法。

原文摘要 · Abstract (English)

Diffusion model alignment aims to bridge the gap between generated outputs and human preferences by enhancing both semantic consistency with textual prompts and overall visual quality. Existing alignment methods face a challenging trade-off: test-time approaches enable input-specific adaptability but introduce significant computational overhead and tend to under-optimize, while fine-tuning approaches risk reward over-optimization and loss of generation diversity. To bridge this gap, we propose HyperAlign, a framework that trains a hypernetwork for efficient and effective test-time alignment. Instead of modifying latent states directly, HyperAlign dynamically generates input-and-state-conditioned low-rank adaptation weights to modulate the denoising trajectory toward target rewards. We introduce multiple HyperAlign variants of varying granularity to balance alignment quality and computational efficiency. The hypernetwork is optimized with a reward objective regularized by preference data to mitigate reward hacking. We evaluate HyperAlign across multiple generative paradigms, including Stable Diffusion and FLUX, where it significantly outperforms existing alignment methods in semantic consistency and visual quality.

扩散模型测试时对齐超网络生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。