用ChatGPT优化文本提示,让AI画画更精准
GPTDrawer: Enhancing Visual Synthesis through ChatGPT
- 用关键词提取+语义分析迭代优化提示词
- 通过余弦相似度判断图文匹配度,直到达标
- 适合需要高精度图文生成的创意设计场景
在人工智能图像生成领域,如何让生成结果精准响应文本提示仍是关键挑战。本文提出GPTDrawer,一个融合ChatGPT自然语言处理与Stable Diffusion图像生成的创新流程。该方法通过关键词提取、语义分析和图文一致性评估,迭代优化输入提示,并利用余弦相似度衡量生成图像与提示的语义对齐程度,直至达到预设阈值。实验表明,该系统显著提升了用户指定提示下图像的生成保真度,展现出对复杂语义结构的准确理解与可视化能力。本研究为创意艺术、设计自动化等应用提供了新的技术范式,推动了AI辅助创作的边界。
原文摘要 · Abstract (English)
In the burgeoning field of AI-driven image generation, the quest for precision and relevance in response to textual prompts remains paramount. This paper introduces GPTDrawer, an innovative pipeline that leverages the generative prowess of GPT-based models to enhance the visual synthesis process. Our methodology employs a novel algorithm that iteratively refines input prompts using keyword extraction, semantic analysis, and image-text congruence evaluation. By integrating ChatGPT for natural language processing and Stable Diffusion for image generation, GPTDrawer produces a batch of images that undergo successive refinement cycles, guided by cosine similarity metrics until a threshold of semantic alignment is attained. The results demonstrate a marked improvement in the fidelity of images generated in accordance with user-defined prompts, showcasing the system's ability to interpret and visualize complex semantic constructs. The implications of this work extend to various applications, from creative arts to design automation, setting a new benchmark for AI-assisted creative processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。