用可逆网络+扩散模型,精准还原被遮挡图像细节。
Reversible Unfolding Network for Concealed Visual Perception with Generative Refinement
- 将感知任务拆解为多阶段可逆网络,跨掩码与彩色域建模。
- 引入针对性扩散机制,仅在不确定区域精细修复,降低误检率。
- 适合应对真实场景中模糊、遮挡等复杂退化问题的研究者。
现有隐含视觉感知(CVP)方法多在掩码域采用可逆策略以降低不确定性,但对彩色域潜力挖掘不足。为此,我们提出可逆展开网络与生成精修结合的RUN++。该方法将CVP建模为数学优化问题,将其迭代求解过程展开为多阶段深度网络,实现掩码与RGB域的统一可逆建模,并利用扩散模型解决不确定性。网络每阶段包含三个专有模块:核心物体区域提取(CORE)在掩码域应用可逆建模;上下文感知区域增强(CARE)将此原则拓展至RGB域,提升前景背景分离;基于噪声增强的微调迭代(FINE)模块引入定向伯努利扩散模型,仅对分割掩码中不确定区域进行精修,利用扩散生成能力恢复细节,避免全图计算开销。该设计使扩散模型聚焦于模糊区域,显著减少误检与漏检。此外,我们提出新范式,构建在真实退化下仍鲁棒的CVP系统,并扩展为更广泛的双层优化框架。
原文摘要 · Abstract (English)
Existing methods for concealed visual perception (CVP) often leverage reversible strategies to decrease uncertainty, yet these are typically confined to the mask domain, leaving the potential of the RGB domain underexplored. To address this, we propose a reversible unfolding network with generative refinement, termed RUN++. Specifically, RUN++ first formulates the CVP task as a mathematical optimization problem and unfolds the iterative solution into a multi-stage deep network. This approach provides a principled way to apply reversible modeling across both mask and RGB domains while leveraging a diffusion model to resolve the resulting uncertainty. Each stage of the network integrates three purpose-driven modules: a Concealed Object Region Extraction (CORE) module applies reversible modeling to the mask domain to identify core object regions; a Context-Aware Region Enhancement (CARE) module extends this principle to the RGB domain to foster better foreground-background separation; and a Finetuning Iteration via Noise-based Enhancement (FINE) module provides a final refinement. The FINE module introduces a targeted Bernoulli diffusion model that refines only the uncertain regions of the segmentation mask, harnessing the generative power of diffusion for fine-detail restoration without the prohibitive computational cost of a full-image process. This unique synergy, where the unfolding network provides a strong uncertainty prior for the diffusion model, allows RUN++ to efficiently direct its focus toward ambiguous areas, significantly mitigating false positives and negatives. Furthermore, we introduce a new paradigm for building robust CVP systems that remain effective under real-world degradations and extend this concept into a broader bi-level optimization framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。