统一去除镜头眩光、雾霾等软性图像退化,一个模型搞定
UniSER: A Foundation Model for Unified Soft Effects Removal
- 基于半透明遮挡的共性,构建统一修复框架
- 380万对数据训练,实现野外真实场景高保真修复
- 适合需要多退化类型联合修复的图像处理场景
数字图像常受镜头眩光、雾霾、阴影和反射等软性退化影响,虽底层像素仍可见但美感下降。现有方法各自为战,依赖高度专业化的模型,难以扩展且未能挖掘各类退化问题的共同本质。尽管近期大模型(如GPT-4o、Flux Kontext、Nano Banana)具备强大文本编辑能力,但在细粒度修复任务中仍依赖复杂提示,难以稳定去除非刚性退化并保持场景一致性。本文基于软性退化共性——半透明遮挡,提出统一基础模型UniSER,可在单一框架内处理多种由软性效应引起的退化。方法核心是构建包含380万对样本的海量数据集,涵盖新型物理合理数据以填补公开基准空白,并设计定制化训练流程,微调扩散变压器(Diffusion Transformer),从多样化数据中学习鲁棒修复先验,集成细粒度掩码与强度控制。该协同策略使UniSER显著超越专用与通用模型,在真实场景下实现稳健、高保真的修复效果。
原文摘要 · Abstract (English)
Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underlying pixels remain partially visible. The prevailing works address these degradations in isolation, developing highly specialized, specialist models that lack scalability and fail to exploit the shared underlying essences of these restoration problems. Meanwhile, although recent large-scale generalist models (e.g., GPT-4o, Flux Kontext, Nano Banana) offer powerful text-driven editing capabilities, they heavily rely on detailed prompts and often fail to achieve robust removal on such fine-grained tasks while preserving the scene's identity. Leveraging the common essence of soft effects, i.e., semi-transparent occlusions, we introduce a foundational versatile model UniSER, capable of addressing diverse degradations caused by soft effects within a single framework. Our methodology centers on curating a massive 3.8M-pair dataset to ensure robustness and generalization, which includes novel, physically-plausible data to fill critical gaps in public benchmarks, and a tailored training pipeline that fine-tunes a Diffusion Transformer to learn robust restoration priors from this diverse data, integrating fine-grained mask and strength controls. This synergistic approach allows UniSER to significantly outperform both specialist and generalist models, achieving robust, high-fidelity restoration in the wild.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。