arXiv:2602.14464cs.CVcs.AI2026-02被引 1

无需训练,实现像素级语义对齐的精细风格迁移。

CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style Transfer

  • 利用扩散模型中间特征构建像素级对应关系图。
  • 通过循环一致性约束保持物体结构与细节不变。
  • 无需额外训练,效果超越需标注的方法。

在保留相似物体间语义对应关系的前提下进行图像风格迁移,仍是计算机视觉中的核心挑战。现有方法多在全局层面操作,忽视了局部乃至像素级别的语义对齐。为此,我们提出 CoCoDiff,一种无需训练、低成本的风格迁移框架,利用预训练的潜在扩散模型实现细粒度、语义一致的风格化。我们发现生成式扩散模型中的对应线索未被充分挖掘,且语义匹配区域间的内容一致性常被忽略。CoCoDiff 引入像素级语义对应模块,从扩散过程中间特征中提取信息,构建内容图与风格图之间的密集对齐映射;同时,循环一致性模块在迭代中强制结构与感知对齐,实现物体及区域级别的风格迁移,有效保留几何形状与细节。尽管无需额外训练或监督,CoCoDiff 在视觉质量与定量指标上均达到领先水平,优于依赖额外训练或标注的方法。

原文摘要 · Abstract (English)

Transferring visual style between images while preserving semantic correspondence between similar objects remains a central challenge in computer vision. While existing methods have made great strides, most of them operate at global level but overlook region-wise and even pixel-wise semantic correspondence. To address this, we propose CoCoDiff, a novel training-free and low-cost style transfer framework that leverages pretrained latent diffusion models to achieve fine-grained, semantically consistent stylization. We identify that correspondence cues within generative diffusion models are under-explored and that content consistency across semantically matched regions is often neglected. CoCoDiff introduces a pixel-wise semantic correspondence module that mines intermediate diffusion features to construct a dense alignment map between content and style images. Furthermore, a cycle-consistency module then enforces structural and perceptual alignment across iterations, yielding object and region level stylization that preserves geometry and detail. Despite requiring no additional training or supervision, CoCoDiff delivers state-of-the-art visual quality and strong quantitative results, outperforming methods that rely on extra training or annotations.

风格迁移扩散模型语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。