arXiv:2511.18164cs.CVcs.AI2025-11被引 6

提出嵌套展开网络,实现真实场景隐匿物体分割的高鲁棒性

Nested Unfolding Network for Real-World Concealed Object Segmentation

  • 采用双重展开结构,分离图像修复与分割任务
  • 动态感知退化类型,提升在复杂真实场景下的分割精度
  • 适合需要高鲁棒性的真实世界隐匿物检测应用

深度展开网络(DUN)通过迭代前景-背景分离推进了隐匿物体分割(COS)。然而现有基于DUN的方法(RUN)将背景估计与图像恢复耦合,目标冲突且需预设退化类型,不适用于真实场景。为此,我们提出嵌套展开网络(NUN),一种统一的实时世界COS框架。NUN采用DUN-in-DUN设计,在分割导向的展开网络(SODUN)每一阶段内嵌入抗退化展开网络(DeRUN),实现修复与分割解耦并支持相互优化。受视觉语言模型(VLM)引导,DeRUN动态推断退化语义并恢复高质量图像,无需显式先验;而SODUN执行可逆估计以精炼前景与背景。利用展开的多阶段特性,NUN通过图像质量评估选择最佳DeRUN输出用于后续阶段,自然引入自一致性损失以增强鲁棒性。大量实验表明,NUN在干净与退化基准上均达到领先性能。代码将开源。

原文摘要 · Abstract (English)

Deep unfolding networks (DUNs) have recently advanced concealed object segmentation (COS) by modeling segmentation as iterative foreground-background separation. However, existing DUN-based methods (RUN) inherently couple background estimation with image restoration, leading to conflicting objectives and requiring pre-defined degradation types, which are unrealistic in real-world scenarios. To address this, we propose the nested unfolding network (NUN), a unified framework for real-world COS. NUN adopts a DUN-in-DUN design, embedding a degradation-resistant unfolding network (DeRUN) within each stage of a segmentation-oriented unfolding network (SODUN). This design decouples restoration from segmentation while allowing mutual refinement. Guided by a vision-language model (VLM), DeRUN dynamically infers degradation semantics and restores high-quality images without explicit priors, whereas SODUN performs reversible estimation to refine foreground and background. Leveraging the multi-stage nature of unfolding, NUN employs image-quality assessment to select the best DeRUN outputs for subsequent stages, naturally introducing a self-consistency loss that enhances robustness. Extensive experiments show that NUN achieves a leading place on both clean and degraded benchmarks. Code will be released.

隐匿物体分割展开网络多模态鲁棒分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。