arXiv:2606.16241cs.CV2026-06

用新框架快速生成高质量视觉文字谜题图像

Structure-Semantic Co-optimized Latent Diffusion Model for Fast Visual Anagram Synthesis

论文配图:Structure-Semantic Co-optimized Latent Diffusion Model for Fast Visual Anagram Synthesis
图 1 · 摘自论文原文
  • 通过结构语义协同优化提升图像一致性
  • 生成分辨率更高、视觉更和谐的图像,速度更快
  • 适合数字艺术创作与高效生成任务

视觉文字谜题是一种艺术形式,同一图像经翻转或旋转后可呈现不同概念含义。现有基于预训练文本到图像扩散模型的方法在计算效率、美学质量及语义保真度方面仍存在不足。本文提出结构-语义协同优化(S2CO)框架,将像素级扩散模型的并行去噪算法迁移至对抗蒸馏的潜空间模型中,显著降低计算开销。核心创新包括:(1) 空文本结构对齐优化;(2) 语义增强优化;(3) 注意力引导噪声融合。所提方法S2CO-Anagram在保持高速推理的同时,生成更高分辨率、更强视觉协调性与语义忠实性的视觉文字谜题图像,优于现有最先进方法。代码将公开。

原文摘要 · Abstract (English)

Visual anagram is an intriguing form of art creation wherein a single image presents different conceptual interpretations under transformations such as flipping or rotation. Recent work has achieved visual anagram synthesis by leveraging pretrained text-to-image (T2I) diffusion models, yet still suffers from several key limitations including computational inefficiency, suboptimal aesthetic quality, and weak semantic fidelity and expressiveness. This work focuses on generating visual anagrams with substantially improved visual quality at minimal computational cost, thereby advancing intelligent creation of illusionary digital art. To increase image resolution while reducing time overhead, we adapt the cutting-edge parallel denoising algorithm from pixel-based T2I model to the adversarially distilled latent-based one, and accordingly propose a structure-semantic co-optimization (S2CO) framework to counteract the consequent visual degradation. As the core of our approach, S2CO framework comprises three key innovations: (\romannumeral1) null-text structure alignment optimization; (\romannumeral2) semantic enhancement optimization; (\romannumeral3) attention-guided noise fusion. Building upon these components, our method dubbed \textbf{S2CO-Anagram} is able to generate higher-resolution anagram images with noticeably superior visual harmony and semantic faithfulness than related SOTA approaches, all while achieving substantially faster inference speed. Code will be publicly available.

图像生成扩散模型视觉艺术高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。