用视觉提示解耦修复,让模糊图像恢复更真实
VICR: Visual In-Context Restoration for Real-World Image Super-Resolution

- 将超分辨率视为图像补全,分离局部结构与全局语义线索
- 在多个真实场景数据集上超越现有方法,仅用127M参数
- 推理时动态优化语义提示,适合处理严重退化的图像
真实世界图像超分辨率(Real-ISR)需在保持退化图像结构真实性与生成合理细节之间取得平衡。现有生成式方法常依赖纠缠的条件机制,导致结构偏移或语义不一致。为此,本文提出视觉上下文修复(VICR),一种基于扩散变压器(DiT)的框架,将Real-ISR建模为图像补全任务。具体地,引入解耦的视觉先验注入机制,从低质量(LQ)图像中提取局部与全局线索:局部线索用于恢复结构并支持高频细节生成,全局线索则引导整体生成并增强语义一致性。对于严重退化导致的模糊区域,VICR在推理时使用一个代理模块,结合LQ输入的视觉证据动态优化语义提示,同时保持模型参数不变。实验表明,VICR在多个Real-ISR基准上达到当前最优性能,仅需127M可训练参数。
原文摘要 · Abstract (English)
Real-world image super-resolution (Real-ISR) requires balancing structural fidelity to degraded observations with realistic detail synthesis. However, existing generative Real-ISR methods often rely on entangled conditioning mechanisms, leading to structural drift or semantically inconsistent details. To address this issue, we propose Visual In-Context Restoration (VICR), a Diffusion Transformer (DiT)-based framework that formulates Real-ISR as image completion. Specifically, we introduce a decoupled visual prior injection mechanism that derives local and global cues from the low-quality (LQ) image: local cues help recover image structures and support high-frequency detail synthesis, while global cues guide overall generation and promote semantic consistency. For ambiguous regions under severe degradation, VICR employs an inference-time agent to refine semantic prompts using visual evidence from the LQ input while keeping model parameters fixed. Experiments show that VICR achieves state-of-the-art performance across multiple Real-ISR benchmarks with only 127M trainable parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。