用少量通道调控提升冻结扩散模型的超分辨率质量
SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution

- 仅优化选定的少数通道,通过输入条件控制实现轻量调优
- 在DIV2K等数据集上显著提升图像保真度与视觉质量
- 适合希望不微调主模型却提升生成效果的研究者
真实世界图像超分辨率(SR)越来越多依赖于扩散Transformer(DiT)骨干网络,其内部激活主要由少数大通道主导。然而,当前提升感知质量仍需微调网络或添加额外适配器,导致这种结构化激活空间未被充分探索。本文研究是否可利用主导通道作为冻结DiT-SR模型的紧凑适配接口。首先分析预训练SR骨干中主导通道的行为,并通过受控干预证明其对重建质量有显著影响。基于此,提出SPARK:一种轻量级输入条件控制器,仅对选定通道预测有界逐通道仿射变换,同时保持SR骨干与VAE完全冻结。主导通道通过在线激活排序识别,仅优化一个基于低分辨率VAE隐变量的微型预测器。在三个基于DiT的SR骨干及DIV2K、RealSR、DRealSR数据集上的实验表明,每流每块仅调节8个通道,即可持续提升保真度与感知质量。受控对比显示,这些增益无法仅由参数预算或访问选定通道解释。
原文摘要 · Abstract (English)
Real-world image super-resolution (SR) increasingly relies on Diffusion Transformer (DiT) backbones, whose internal activations can be dominated by a small number of massive channels. Yet improving perceptual quality in these models still typically requires fine-tuning the network or attaching additional adapters, leaving this structured activation space largely unexplored for adaptation. We investigate whether dominant channels can instead serve as a compact adaptation interface for frozen DiT-based SR models. We first characterize their behavior in pretrained SR backbones and show through controlled interventions that they strongly affect reconstruction quality. Building on this observation, we introduce SPARK, a lightweight input-conditioned controller that predicts bounded per-channel affine transformations for only the selected channels, while keeping the SR backbone and VAE frozen. Dominant channels are identified through an online activation-ranking procedure, and only a small predictor conditioned on the low-resolution VAE latent is optimized. Experiments on three DiT-based SR backbones across DIV2K, RealSR, and DRealSR show consistent gains in both fidelity and perceptual quality while modulating only eight channels per stream and block. Controlled comparisons further show that these gains cannot be explained by parameter budget or access to the selected channels alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。