arXiv:2604.04780cs.CV2026-04被引 3

让统一模型在模糊图像上先生成再理解,提升真实场景下的多模态推理能力。

CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models

  • 通过三步框架让模型学会先生成后推理的思维模式
  • 在六项基准上对劣质图像的鲁棒性显著提升,且保持清晰图像性能
  • 适合需要处理真实世界退化图像的多模态应用开发者

由模糊、噪声、压缩和光照不足引起的图像退化严重损害了真实场景中多模态理解的效果。统一多模态模型将理解与生成整合于单一架构,天然适合应对该挑战,因其生成路径可重建退化所破坏的细粒度视觉结构。然而,现有模型未能利用自身生成能力处理退化输入。我们发现这一断层源于两个因素:训练未要求模型在推理时调用生成,且标准的解码-重编码路径不支持有效联合优化。本文提出CLEAR框架,通过三个渐进步骤实现双能力连接:(1) 在退化感知数据集上进行监督微调,建立‘生成-回答’推理模式;(2) 引入潜在表示桥,以直接可优化连接替代解码-重编码路径;(3) 采用交错式GRPO强化学习方法,在答案正确性奖励下联合优化文本推理与视觉生成。我们构建MMD-Bench,涵盖六项标准多模态基准上的三种退化强度。实验表明,CLEAR显著提升退化输入下的鲁棒性,同时保持清晰图像性能。分析进一步揭示,移除像素级重建监督后,中间视觉状态具有更高感知质量,暗示任务驱动优化与视觉质量天然一致。

原文摘要 · Abstract (English)

Image degradation from blur, noise, compression, and poor illumination severely undermines multimodal understanding in real-world settings. Unified multimodal models that combine understanding and generation within a single architecture are a natural fit for this challenge, as their generative pathway can model the fine-grained visual structure that degradation destroys. Yet these models fail to leverage their own generative capacity on degraded inputs. We trace this disconnect to two compounding factors: existing training regimes never ask the model to invoke generation during reasoning, and the standard decode-reencode pathway does not support effective joint optimization. We present CLEAR, a framework that connects the two capabilities through three progressive steps: (1) supervised fine-tuning on a degradation-aware dataset to establish the generate-then-answer reasoning pattern; (2) a Latent Representation Bridge that replaces the decode-reencode detour with a direct, optimizable connection between generation and reasoning; (3) Interleaved GRPO, a reinforcement learning method that jointly optimizes text reasoning and visual generation under answer-correctness rewards. We construct MMD-Bench, covering three degradation severity levels across six standard multimodal benchmarks. Experiments show that CLEAR substantially improves robustness on degraded inputs while preserving clean-image performance. Our analysis further reveals that removing pixel-level reconstruction supervision leads to intermediate visual states with higher perceptual quality, suggesting that task-driven optimization and visual quality are naturally aligned.

多模态图像修复生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。