arXiv:2505.17618cs.CVcs.AI2025-05被引 36

用进化搜索提升图像视频生成质量,推理时高效扩增算力。

Scaling Image and Video Generation via Test-Time Evolutionary Search

  • 将生成过程视为进化搜索,迭代优化去噪路径。
  • 在图像和视频生成上均超越现有方法,样本多样性更高。
  • 无需额外训练,通用性强,适合各类扩散与流模型。

随着预训练阶段计算成本(数据与参数)持续上升,测试时扩展(TTS)作为一种在推理阶段分配额外计算以提升生成模型性能的策略,正成为重要方向。尽管TTS在语言任务中表现显著,但对基于扩散或流模型的图像与视频生成模型的测试时扩展行为仍缺乏深入理解。现有方法存在领域特定、可扩展性差或奖励过优化导致样本多样性下降等问题。本文提出一种名为EvoSearch的新方法,通过将测试时扩展重构为进化搜索问题,利用生物进化原理高效探索并优化去噪轨迹。通过针对随机微分方程去噪过程设计的选择与变异机制,EvoSearch在保持种群多样性的同时,迭代生成更高质量的样本。在多种扩散与流架构的图像与视频生成任务上广泛评估表明,该方法始终优于现有方法,在多样性与泛化性方面表现优异,且无需额外训练或模型扩展。项目主页:https://tinnerhrhe.github.io/evosearch。

原文摘要 · Abstract (English)

As the marginal cost of scaling computation (data and parameters) during model pre-training continues to increase substantially, test-time scaling (TTS) has emerged as a promising direction for improving generative model performance by allocating additional computation at inference time. While TTS has demonstrated significant success across multiple language tasks, there remains a notable gap in understanding the test-time scaling behaviors of image and video generative models (diffusion-based or flow-based models). Although recent works have initiated exploration into inference-time strategies for vision tasks, these approaches face critical limitations: being constrained to task-specific domains, exhibiting poor scalability, or falling into reward over-optimization that sacrifices sample diversity. In this paper, we propose \textbf{Evo}lutionary \textbf{Search} (EvoSearch), a novel, generalist, and efficient TTS method that effectively enhances the scalability of both image and video generation across diffusion and flow models, without requiring additional training or model expansion. EvoSearch reformulates test-time scaling for diffusion and flow models as an evolutionary search problem, leveraging principles from biological evolution to efficiently explore and refine the denoising trajectory. By incorporating carefully designed selection and mutation mechanisms tailored to the stochastic differential equation denoising process, EvoSearch iteratively generates higher-quality offspring while preserving population diversity. Through extensive evaluation across both diffusion and flow architectures for image and video generation tasks, we demonstrate that our method consistently outperforms existing approaches, achieves higher diversity, and shows strong generalizability to unseen evaluation metrics. Our project is available at the website https://tinnerhrhe.github.io/evosearch.

图像生成视频生成扩散模型进化搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。