提出新方法实现零样本图像渐变,结果更平滑、语义更连贯。
CHIMERA: Adaptive Cache Injection and Semantic Anchor Prompting for Zero-shot Image Morphing with Morphing-oriented Metrics
- 通过自适应缓存注入复用多尺度特征,稳定生成过程。
- 引入语义锚点提示,使中间图像语义一致性提升显著。
- 适合需要高质量渐变且不想重训练的开发者使用。
基于扩散模型的图像渐变方法通常在反演潜在表示时复用有限条件信号,导致异构起点对之间生成不稳定。主要问题包括:(i) 特征复用常为部分或非自适应,造成结构突变或过度平滑;(ii) 文本条件独立获取后插值,易引入语义不兼容。本文提出CHIMERA,一种新型零样本扩散渐变框架,通过反演引导的去噪与互补特征复用、文本条件优化解决上述问题。自适应缓存注入(ACI)在DDIM反演中缓存更广范围的多尺度扩散特征,并以层和时间步感知调度重新注入,增强去噪稳定性与渐进融合能力。语义锚点提示(SAP)利用视觉语言模型生成共享锚点提示与锚定条件端点提示,并将锚点注入交叉注意力,提升中间阶段语义连贯性。最后,提出面向渐变任务的全局-局部一致性评分(GLCS),联合衡量全局领域协调性与局部过渡平滑性。大量实验与用户研究显示,相比现有方法,CHIMERA生成的渐变结果更平滑、语义更一致,同时保持高效性,适用于多种扩散骨干网络且无需再训练。
原文摘要 · Abstract (English)
Recent diffusion-based image morphing methods typically interpolate inverted latents and reuse limited conditioning signals, which often yields unstable intermediates for heterogeneous endpoint pairs. In particular, (i) feature reuse is usually partial or non-adaptive, leading to abrupt structural changes or over-smoothing, and (ii) text conditions are commonly obtained independently per endpoint and then interpolated, which can introduce incompatible semantics. We present CHIMERA, a novel zero-shot diffusion morphing framework that addresses both issues via inversion-guided denoising with complementary feature reuse and text conditioning. Adaptive Cache Injection (ACI) caches a broader set of multi-scale diffusion features beyond Key-Value-only reuse during DDIM inversion, and re-injects them with layer- and timestep-aware scheduling to stabilize denoising and enable gradual fusion. Semantic Anchor Prompting (SAP) uses a VLM to generate a shared anchor-prompt and anchor-conditioned endpoint prompts, and injects the anchor into cross-attention to improve intermediate semantic coherence. Finally, we propose Global-Local Consistency Score (GLCS), a morphing-oriented metric that jointly captures global domain harmonization and local transition smoothness. Extensive experiments and a user study show that CHIMERA produces smoother and more semantically consistent morphing results than prior methods, while remaining efficient and applicable across diverse diffusion backbones without retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。