用多任务学习生成视觉错觉图,解决概念分离和主导问题。
Diffusion-based Visual Anagram as Multi-task Learning
- 将不同视角视为任务,统一优化去噪轨迹。
- 提出抗分离与噪声平衡策略,提升多视角一致性。
- 适合图像生成、视觉幻觉研究者参考。
视觉安格拉姆是通过翻转或旋转等变换改变外观的图像。随着扩散模型的发展,可通过在反向去噪过程中平均多个视角的噪声来生成此类视觉错觉。然而我们发现该方法存在两种关键失效模式:(i) 概念分离,即不同视角中的概念独立生成,无法构成真正的安格拉姆;(ii) 概念主导,某些概念过度压制其他概念。本文将视觉安格拉姆生成问题建模为多任务学习,将不同视角提示视为不同任务,并设计出跨任务同步对齐的去噪轨迹。核心包含两项新方法:(i) 抗分离优化策略,促进不同概念间交叉注意力图的重叠;(ii) 噪声向量平衡方法,自适应调节各任务影响。此外,我们发现直接平均噪声预测性能不佳,因统计特性未被保留,因而提出噪声方差校正方法。大量定性与定量实验表明,本方法能有效生成涵盖多样概念的视觉安格拉姆。
原文摘要 · Abstract (English)
Visual anagrams are images that change appearance upon transformation, like flipping or rotation. With the advent of diffusion models, generating such optical illusions can be achieved by averaging noise across multiple views during the reverse denoising process. However, we observe two critical failure modes in this approach: (i) concept segregation, where concepts in different views are independently generated, which can not be considered a true anagram, and (ii) concept domination, where certain concepts overpower others. In this work, we cast the visual anagram generation problem in a multi-task learning setting, where different viewpoint prompts are analogous to different tasks,and derive denoising trajectories that align well across tasks simultaneously. At the core of our designed framework are two newly introduced techniques, where (i) an anti-segregation optimization strategy that promotes overlap in cross-attention maps between different concepts, and (ii) a noise vector balancing method that adaptively adjusts the influence of different tasks. Additionally, we observe that directly averaging noise predictions yields suboptimal performance because statistical properties may not be preserved, prompting us to derive a noise variance rectification method. Extensive qualitative and quantitative experiments demonstrate our method's superior ability to generate visual anagrams spanning diverse concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。