arXiv:2509.23605cs.CV2025-09被引 1

用扩散模型融合两图,生成无偏且连贯的新物体

VMDiff: Visual Mixing Diffusion for Limitless Cross-Object Synthesis

  • 在噪声和潜在空间同时融合两图,实现结构感知合成
  • 780组概念对测试中,视觉质量与语义一致性超越基线
  • 自动调节参数,解决主次对象失衡问题,适合创意设计

通过融合多个来源的视觉线索生成新图像是图像到图像生成中的基础但未充分探索的问题,广泛应用于艺术创作、虚拟现实和视觉媒体。现有方法常面临两个关键挑战:共存生成(多个物体简单并置而未真正整合)和偏差生成(因语义不平衡导致一个物体主导输出)。为此,我们提出视觉混合扩散(VMDiff),一种基于扩散的简单而有效的框架,通过在噪声和潜在空间同时融合两张输入图像,生成单一连贯物体。该方法包括:(1) 混合采样过程,结合引导去噪、反演与球面插值,可调参数实现结构感知融合,缓解共存生成;(2) 高效自适应调整模块,引入新颖的基于相似性的评分机制,自动且自适应地搜索最优参数,对抗语义偏差。在包含780个概念对的定制基准上实验表明,该方法在视觉质量、语义一致性和人类评估的创造力方面均优于强基线。

原文摘要 · Abstract (English)

Creating novel images by fusing visual cues from multiple sources is a fundamental yet underexplored problem in image-to-image generation, with broad applications in artistic creation, virtual reality and visual media. Existing methods often face two key challenges: coexistent generation, where multiple objects are simply juxtaposed without true integration, and bias generation, where one object dominates the output due to semantic imbalance. To address these issues, we propose Visual Mixing Diffusion (VMDiff), a simple yet effective diffusion-based framework that synthesizes a single, coherent object by integrating two input images at both noise and latent levels. Our approach comprises: (1) a hybrid sampling process that combines guided denoising, inversion, and spherical interpolation with adjustable parameters to achieve structure-aware fusion, mitigating coexistent generation; and (2) an efficient adaptive adjustment module, which introduces a novel similarity-based score to automatically and adaptively search for optimal parameters, countering semantic bias. Experiments on a curated benchmark of 780 concept pairs demonstrate that our method outperforms strong baselines in visual quality, semantic consistency, and human-rated creativity.

图像生成扩散模型跨对象融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。