无需文字提示,用双路信号引导修复真实图像,效果超越主流方法。
LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer
- 通过输入与轻度修复图双信号注入,分别锚定结构与抑制伪影。
- 在合成与真实数据上均超越开源及商业模型,峰值性能达42.3dB。
- 不依赖文本或视觉语言模型,适合实际场景中无标注图像修复。
图像恢复(IR)旨在修复由未知混合因素导致退化的图像,同时保持语义一致性。现有判别式恢复器和基于UNet的扩散先验常导致过度平滑、幻觉或漂移。本文提出LucidFlux,一种无需图像描述的端到端图像恢复框架,适配大规模扩散变换器(Flux.1)。其引入轻量级双分支条件模块,分别从退化输入和轻度修复代理中注入信号,以分别锚定几何结构并抑制伪影。进一步设计时间步与层自适应调制策略,实现从粗到细、上下文感知的更新,保护全局结构的同时恢复纹理细节。为避免文本提示或视觉-语言模型(VLM)生成描述带来的延迟与不稳定性,采用从代理图像提取的SigLIP特征实现无提示语义对齐。通过可扩展的数据清洗流程,筛选出富含结构信息的大规模训练数据。在合成与真实场景基准测试中,LucidFlux持续优于强开源与商业基线,消融实验验证各组件必要性。结果表明,对于大尺度扩散变换器,何时、何地、如何进行条件注入,是实现鲁棒、无提示图像恢复的关键控制杠杆。
原文摘要 · Abstract (English)
Image restoration (IR) aims to recover images degraded by unknown mixtures while preserving semanticsconditions under which discriminative restorers and UNet-based diffusion priors often oversmooth, hallucinate, or drift. We present LucidFlux, a caption-free IR framework that adapts a large diffusion transformer (Flux.1) without image captions. Our LucidFlux introduces a lightweight dual-branch conditioner that injects signals from the degraded input and a lightly restored proxy to respectively anchor geometry and suppress artifacts. Then, a timestep- and layer-adaptive modulation schedule is designed to route these cues across the backbones hierarchy, in order to yield coarse-to-fine and context-aware updates that protect the global structure while recovering texture. After that, to avoid the latency and instability of text prompts or Vision-Language Model (VLM) captions, we enforce caption-free semantic alignment via SigLIP features extracted from the proxy. A scalable curation pipeline further filters large-scale data for structure-rich supervision. Across synthetic and in-the-wild benchmarks, our LucidFlux consistently outperforms strong open-source and commercial baselines, and ablation studies verify the necessity of each component. LucidFlux shows that, for large DiTs, when, where, and what to condition onrather than adding parameters or relying on text promptsis the governing lever for robust and caption-free image restoration in the wild.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。