用测试时缩放让大模型高效修复真实图像,效果超前。
Tuning Real-World Image Restoration at Inference: A Test-Time Scaling Paradigm for Flow Matching Models
- 用统一模态融合让多条件指导图像生成
- 测试时通过奖励模型反馈动态优化去噪,性能显著提升
- 无需训练即可在推理阶段扩展模型能力,适合大模型应用
尽管基于扩散模型的真实世界图像修复(Real-IR)已取得显著进展,但如何高效利用超大规模预训练文本到图像(T2I)模型并充分发挥其潜力仍是重大挑战。为此,我们提出ResFlow-Tuner,一个基于先进流匹配模型FLUX.1-dev的图像修复框架,结合统一多模态融合(UMMF)与测试时缩放(TTS),实现前所未有的修复性能。该方法充分利用多模态扩散Transformer(MM-DiT)架构优势,将多模态条件编码为统一序列以引导高质量图像合成。此外,我们引入一种无需训练的测试时缩放范式:在推理过程中,通过奖励模型(RM)的反馈动态调整去噪方向,从而在可控计算开销下获得显著性能提升。大量实验表明,该方法在多个标准基准上达到领先水平。本工作不仅验证了流匹配模型在低层视觉任务中的强大能力,更提出一种适用于大型预训练模型的新颖高效推理时扩展范式。
原文摘要 · Abstract (English)
Although diffusion-based real-world image restoration (Real-IR) has achieved remarkable progress, efficiently leveraging ultra-large-scale pre-trained text-to-image (T2I) models and fully exploiting their potential remain significant challenges. To address this issue, we propose ResFlow-Tuner, an image restoration framework based on the state-of-the-art flow matching model, FLUX.1-dev, which integrates unified multi-modal fusion (UMMF) with test-time scaling (TTS) to achieve unprecedented restoration performance. Our approach fully leverages the advantages of the Multi-Modal Diffusion Transformer (MM-DiT) architecture by encoding multi-modal conditions into a unified sequence that guides the synthesis of high-quality images. Furthermore, we introduce a training-free test-time scaling paradigm tailored for image restoration. During inference, this technique dynamically steers the denoising direction through feedback from a reward model (RM), thereby achieving significant performance gains with controllable computational overhead. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple standard benchmarks. This work not only validates the powerful capabilities of the flow matching model in low-level vision tasks but, more importantly, proposes a novel and efficient inference-time scaling paradigm suitable for large pre-trained models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。