用64个词元模板提升文本生成图像的创意,让设计更高效。
A Creative Agent is Worth a 64-Token Template
- 通过创意分拆训练,将创作理解封装为可复用的词元模板。
- 在建筑、家具等任务中实现3.7倍加速和4.8倍成本降低。
- 适合需要高频创意生成的研究者与设计师使用。
文本到图像(T2I)模型在图像质量和提示遵循方面已有显著提升,但其创意能力仍受限于对离散自然语言提示的依赖。当面对模糊提示如“受黑胶唱片启发的创意摩天楼”时,模型难以推断深层创意意图,导致创意构思与提示设计仍需人工完成。现有基于推理或智能体的方法虽能迭代优化提示,但计算与成本高昂,因每次生成需独立推理,创意无法复用。为此,我们提出CAT框架——创意智能体词元化,通过创意分拆训练,将智能体对“创意”的内在理解封装为可复用的词元模板。给定模糊提示的嵌入表示,该模板可直接拼接至提示中,无需重复推理即可注入创意语义。在建筑、家具设计及自然混合任务上的实验证明,CAT实现了可扩展且高效的创意增强,相比先进T2I模型与创意生成方法,在人类偏好和图文对齐上表现更优,同时带来3.7倍速度提升和4.8倍计算成本下降。
原文摘要 · Abstract (English)
Text-to-image (T2I) models have substantially improved image fidelity and prompt adherence, yet their creativity remains constrained by reliance on discrete natural language prompts. When presented with fuzzy prompts such as ``a creative vinyl record-inspired skyscraper'', these models often fail to infer the underlying creative intent, leaving creative ideation and prompt design largely to human users. Recent reasoning- or agent-driven approaches iteratively augment prompts but incur high computational and monetary costs, as their instance-specific generation makes ``creativity'' costly and non-reusable, requiring repeated queries or reasoning for subsequent generations. To address this, we introduce \textbf{CAT}, a framework for \textbf{C}reative \textbf{A}gent \textbf{T}okenization that encapsulates agents' intrinsic understanding of ``creativity'' through a \textit{Creative Tokenizer}. Given the embeddings of fuzzy prompts, the tokenizer generates a reusable token template that can be directly concatenated with them to inject creative semantics into T2I models without repeated reasoning or prompt augmentation. To enable this, the tokenizer is trained via creative semantic disentanglement, leveraging relations among partially overlapping concept pairs to capture the agent's latent creative representations. Extensive experiments on \textbf{\textit{Architecture Design}}, \textbf{\textit{Furniture Design}}, and \textbf{\textit{Nature Mixture}} tasks demonstrate that CAT provides a scalable and effective paradigm for enhancing creativity in T2I generation, achieving a $3.7\times$ speedup and a $4.8\times$ reduction in computational cost, while producing images with superior human preference and text-image alignment compared to state-of-the-art T2I models and creative generation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。