无需重训,测试时优化图像修复结果的偏好对齐方法
Test-Time Preference Optimization for Image Restoration
- 通过扩散反演生成候选修复图像,实现在线偏好数据生成
- 利用自动化或人工反馈筛选偏好图像,指导修复过程优化
- 兼容任意模型架构,适配多种图像修复任务且无需重新训练
图像修复(IR)模型通常使用L1或LPIPS损失进行训练。为应对未知退化情况,零样本修复方法也被提出。然而,现有预训练和零样本修复方法常无法与人类偏好对齐,导致修复结果不被青睐。这凸显了提升修复质量并灵活适应不同任务或模型主干的迫切需求,且无需模型重训,理想情况下也无需耗时的人工偏好数据收集。本文首次提出测试时偏好优化(TTPO)范式,可增强感知质量、在线生成偏好数据,并兼容任意图像修复模型主干。具体设计了一个无训练、三阶段流程:(i) 基于初始修复图像,通过扩散反演与去噪生成候选偏好图像;(ii) 使用自动偏好对齐度量或人工反馈筛选偏好与非偏好图像;(iii) 将选定图像作为奖励信号,引导扩散去噪过程,优化修复结果以更好匹配人类偏好。在多种图像修复任务与模型上的大量实验表明,该方法有效且灵活。
原文摘要 · Abstract (English)
Image restoration (IR) models are typically trained to recover high-quality images using L1 or LPIPS loss. To handle diverse unknown degradations, zero-shot IR methods have also been introduced. However, existing pre-trained and zero-shot IR approaches often fail to align with human preferences, resulting in restored images that may not be favored. This highlights the critical need to enhance restoration quality and adapt flexibly to various image restoration tasks or backbones without requiring model retraining and ideally without labor-intensive preference data collection. In this paper, we propose the first Test-Time Preference Optimization (TTPO) paradigm for image restoration, which enhances perceptual quality, generates preference data on-the-fly, and is compatible with any IR model backbone. Specifically, we design a training-free, three-stage pipeline: (i) generate candidate preference images online using diffusion inversion and denoising based on the initially restored image; (ii) select preferred and dispreferred images using automated preference-aligned metrics or human feedback; and (iii) use the selected preference images as reward signals to guide the diffusion denoising process, optimizing the restored image to better align with human preferences. Extensive experiments across various image restoration tasks and models demonstrate the effectiveness and flexibility of the proposed pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。