arXiv:2502.05606cs.CV2025-02被引 8

无需训练,通过分步反馈机制实现更自然的图像概念融合。

FreeBlend: Advancing Concept Blending with Staged Feedback-Driven Interpolation Diffusion

  • 分阶段渐进插值,逐步调整融合比例以优化特征整合
  • 反向反馈更新辅助潜在表示,提升整体融合效果
  • 无需训练,适合快速生成语义一致且视觉高质量的融合图像

概念融合是生成模型中一个有前景但研究不足的领域。尽管已有基于嵌入混合和结构草图的潜在空间修改方法,但仍面临语义不兼容、形状与外观差异等问题。本文提出FreeBlend,一种无需训练的高效框架,通过转移图像嵌入作为条件输入,缓解跨模态损失并增强特征细节。该框架采用逐步递增的潜在变量插值策略,动态调整融合比例,实现辅助特征的无缝集成;同时引入反向反馈机制,逆序更新辅助潜在变量,促进全局融合,避免输出僵硬或不自然。大量实验表明,该方法显著提升了融合图像的语义连贯性与视觉质量,生成结果更具说服力与一致性。

原文摘要 · Abstract (English)

Concept blending is a promising yet underexplored area in generative models. While recent approaches, such as embedding mixing and latent modification based on structural sketches, have been proposed, they often suffer from incompatible semantic information and discrepancies in shape and appearance. In this work, we introduce FreeBlend, an effective, training-free framework designed to address these challenges. To mitigate cross-modal loss and enhance feature detail, we leverage transferred image embeddings as conditional inputs. The framework employs a stepwise increasing interpolation strategy between latents, progressively adjusting the blending ratio to seamlessly integrate auxiliary features. Additionally, we introduce a feedback-driven mechanism that updates the auxiliary latents in reverse order, facilitating global blending and preventing rigid or unnatural outputs. Extensive experiments demonstrate that our method significantly improves both the semantic coherence and visual quality of blended images, yielding compelling and coherent results.

概念融合扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。