arXiv:2602.04565cs.CV2026-02

用视觉语言模型精准识别图像退化的类型与物理参数。

Understanding Degradation with Vision Language Model

论文配图:Understanding Degradation with Vision Language Model
图 1 · 摘自论文原文
  • 将退化理解为分层结构预测任务,统一建模类型、参数键与连续值。
  • 在11万对图像数据上训练,恢复精度显著优于现有方法。
  • 可零样本控制生成模型,无需微调即可实现高质量修复。

理解视觉退化是计算机视觉中的关键但具挑战性的问题。尽管近期的视觉语言模型(VLMs)在定性描述方面表现优异,但在理解图像退化背后的参数化物理机制方面仍显不足。本文将退化理解重新定义为一种分层结构化预测任务,需要同时估计退化类型、参数键及其连续物理值。虽然这些子任务处于不同空间,但我们证明它们可在单一自回归下一个标记预测范式下统一,其误差受值空间量化网格约束。基于此,我们提出DU-VLM,一种使用监督微调和强化学习(带结构奖励)训练的多模态思维链模型。此外,我们展示DU-VLM可作为预训练扩散模型的零样本控制器,无需微调生成主干即可实现高保真图像修复。我们还引入了 extbf{DU-110k},一个包含11万对清晰-退化图像对并附有物理标注的大型数据集。大量实验表明,该方法在准确性和鲁棒性上均显著优于通用基线,在未见分布上也表现出良好泛化能力。

原文摘要 · Abstract (English)

Understanding visual degradations is a critical yet challenging problem in computer vision. While recent Vision-Language Models (VLMs) excel at qualitative description, they often fall short in understanding the parametric physics underlying image degradations. In this work, we redefine degradation understanding as a hierarchical structured prediction task, necessitating the concurrent estimation of degradation types, parameter keys, and their continuous physical values. Although these sub-tasks operate in disparate spaces, we prove that they can be unified under one autoregressive next-token prediction paradigm, whose error is bounded by the value-space quantization grid. Building on this insight, we introduce DU-VLM, a multimodal chain-of-thought model trained with supervised fine-tuning and reinforcement learning using structured rewards. Furthermore, we show that DU-VLM can serve as a zero-shot controller for pre-trained diffusion models, enabling high-fidelity image restoration without fine-tuning the generative backbone. We also introduce \textbf{DU-110k}, a large-scale dataset comprising 110,000 clean-degraded pairs with grounded physical annotations. Extensive experiments demonstrate that our approach significantly outperforms generalist baselines in both accuracy and robustness, exhibiting generalization to unseen distributions.

视觉退化多模态扩散模型结构预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。