arXiv:2602.20989cs.CV2026-02中稿 · CVPR

用循环一致训练让扩散模型精准分离图像中的图章与背景

Cycle-Consistent Tuning for Layered Image Decomposition

  • 用轻量LoRA微调扩散模型,结合正反向重建约束提升分离精度
  • 在真实图像上实现图章与表面的高质量解耦,保持视觉一致性
  • 支持自迭代优化,适合复杂交互场景的图像分解任务

真实图像中视觉层的解耦是计算机视觉与图形学中的长期挑战,因各层常存在非线性及全局耦合关系,如阴影、反射和透视畸变。本文提出一种基于大扩散基础模型的上下文图像解耦框架,聚焦于图章-物体解耦任务,旨在将图章与其所在表面准确分离并忠实保留两层特征。方法通过轻量级LoRA微调预训练扩散模型,并引入循环一致训练策略,联合优化解耦与重构模型,强制保证分解后图像与重构图像间的一致性,显著增强复杂交互情况下的鲁棒性。此外,提出渐进式自我提升流程,通过高质模型生成样本迭代扩充训练集以持续优化性能。大量实验表明,该方法可实现精确且连贯的解耦效果,并有效泛化至其他分解类型,具备作为统一图像分层解耦框架的潜力。

原文摘要 · Abstract (English)

Disentangling visual layers in real-world images is a persistent challenge in vision and graphics, as such layers often involve non-linear and globally coupled interactions, including shading, reflection, and perspective distortion. In this work, we present an in-context image decomposition framework that leverages large diffusion foundation models for layered separation. We focus on the challenging case of logo-object decomposition, where the goal is to disentangle a logo from the surface on which it appears while faithfully preserving both layers. Our method fine-tunes a pretrained diffusion model via lightweight LoRA adaptation and introduces a cycle-consistent tuning strategy that jointly trains decomposition and composition models, enforcing reconstruction consistency between decomposed and recomposed images. This bidirectional supervision substantially enhances robustness in cases where the layers exhibit complex interactions. Furthermore, we introduce a progressive self-improving process, which iteratively augments the training set with high-quality model-generated examples to refine performance. Extensive experiments demonstrate that our approach achieves accurate and coherent decompositions and also generalizes effectively across other decomposition types, suggesting its potential as a unified framework for layered image decomposition.

图像解耦扩散模型LoRA循环一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。