用最优控制理论提升扩散模型协同生成的一致性与泛化能力
Variational Test-time Optimization for Diffusion Synchronization

- 基于最优控制构建测试时优化框架,动态协调多条扩散路径
- 在三种跨模态任务上显著优于基线,无需额外训练
- 适用于多种生成场景,为协作生成提供新理论基础
协同生成通过协调多个扩散轨迹来拓展预训练先验的能力,已成为扩展扩散模型应用范围的有力范式。现有方法中,扩散同步通过引入通用引导机制,实现了场景无关的解决方案。然而,当前同步方法仍高度依赖启发式设计,且需针对任务进行定制,限制了其泛化性和性能。本文从最优控制角度推导出同步框架,为扩散同步提供了严谨的数学解释。采样过程中,我们优化控制变量,引导多条轨迹趋向一致解,同时保持与底层扩散先验的接近。该方法完全在测试阶段运行,无需额外训练,结合强大的预训练先验,可广泛适用于多样化的生成场景。我们在三个代表性协同生成任务上验证了方法的一致改进,涵盖多种模态与应用。除性能提升外,本工作建立了协同生成的新范式,为将预训练生成模型推广至新的协作生成场景开辟了有原则的路径。
原文摘要 · Abstract (English)
Collaborative generation, which coordinates multiple diffusion trajectories to extend the capabilities of pretrained priors, has emerged as a powerful paradigm for extending the applicability of diffusion models. Among existing approaches, diffusion synchronization provides a scenario-agnostic solution by introducing general guidance mechanisms. However, current synchronization approaches rely heavily on heuristics and still require task-specific tailoring, which limits their generalizability and performance. In this work, we mathematically derive a synchronization framework based on optimal control, providing a principled explanation of diffusion synchronization. During sampling, we optimize control variables to guide multiple trajectories toward coherent solutions while remaining close to the underlying diffusion prior. Our method operates entirely at test-time without additional training, thereby enabling broad applicability across diverse generation scenarios when combined with strong pretrained priors. We demonstrate consistent improvements over baselines on three representative collaborative generation tasks, covering a wide range of modalities and applications. Beyond performance gains, our work establishes a novel foundation for collaborative generation, opening a principled path toward extending pretrained generative models to new collaborative generation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。