arXiv:2605.19554cs.CV2026-05

让文字生成图像更富创意,突破传统模型的平淡输出。

Self-Creative Text-to-Object Generation using Semantic-Aware Spatial Weighting

论文配图:Self-Creative Text-to-Object Generation using Semantic-Aware Spatial Weighting
图 1 · 摘自论文原文
  • 用可学习的空间加权模块增强图像中心特征,激发新颖性。
  • 双损失机制平衡语义对齐与视觉差异,提升创意与一致性。
  • 适合需要艺术化生成的设计师或内容创作者使用。

在文本到图像生成中注入创造力面临重大挑战,需兼顾视觉新颖性、意外感与艺术价值。现有模型多优化于文本与图像的字面匹配,其噪声预测网络限制生成范围在高概率区域,导致结果缺乏真正创意。为此,我们提出自创意扩散模型(SCDiff),包含两个核心模块:可学习空间加权(LSW)模块与视觉-语义混合损失(VSML)。LSW模块设计参数化Kaiser-Bessel窗口,强化图像中心特征,促进新颖且出人意料的生成。VSML模块引入双重损失:相似性损失确保新图像符合文本描述,多样性损失最大化其与原图的差异,从而提升语义价值与视觉新颖性。大量实验表明,该模型显著提升创意性、语义对齐与视觉连贯性,提供一种简单而强大的创造性物体生成框架。

原文摘要 · Abstract (English)

Instilling creativity in text-to-image (T2I) generation presents a significant challenge, as it requires synthesized images to exhibit not only visual novelty and surprise, but also artistic value. Current T2I models, however, are largely optimized for literal text-image alignment with their data distribution, and their noise prediction networks constrain the generation to high-probability regions, consequently generating outputs that lack authentic creativity. To address this, we propose a Self-Creative Diffusion (SCDiff) model for meaningful T2I generations featuring two core modules: a learnable spatial weighting (LSW) module and a visual-semantic mixing loss (VSML). The LSW module designs a parametric Kaiser-Bessel window to reinforce central image features, fostering novel and surprising generation. The VSML module introduces a dual loss function: a similarity loss constrains that the new images align with its textual description, while a diversity loss maximizes its distinction from the original image, enhancing both semantic value and visual novelty. Extensive experiments demonstrate that our model substantially improves creativity, semantic alignment, and visual coherence, offering a simple yet powerful framework for generating creative objects.

文本生成图像创意生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。