让AI画画时能上网查参考图,解决提示词模糊的问题
IA-T2I: Internet-Augmented Text-to-Image Generation
- 通过联网检索图片,补充提示词中不确定的信息
- 在人类评估中比GPT-4o高30%准确率
- 适合需要实时更新视觉风格的创作场景
当前文本到图像生成模型在提示词所隐含知识不确定的场景下表现不佳。例如,2月发布的模型难以生成4月上映电影的海报,因角色设计和风格尚未确定。为此,我们提出Internet-Augmented text-to-image generation(IA-T2I)框架,通过提供参考图像使模型明确不确定知识。具体包括:主动检索模块判断是否需参考图;分层图像选择模块从图像搜索引擎返回结果中筛选最适配图像;自反思机制持续评估并优化生成图像,确保与提示词一致。为评估性能,我们构建了名为Img-Ref-T2I的数据集,包含三类不确定知识:(1)已知但罕见;(2)未知;(3)模糊。此外,我们精心设计复杂提示引导GPT-4o进行偏好评估,其准确性接近人类。实验表明,本框架在人类评估中优于GPT-4o约30%。
原文摘要 · Abstract (English)
Current text-to-image (T2I) generation models achieve promising results, but they fail on the scenarios where the knowledge implied in the text prompt is uncertain. For example, a T2I model released in February would struggle to generate a suitable poster for a movie premiering in April, because the character designs and styles are uncertain to the model. To solve this problem, we propose an Internet-Augmented text-to-image generation (IA-T2I) framework to compel T2I models clear about such uncertain knowledge by providing them with reference images. Specifically, an active retrieval module is designed to determine whether a reference image is needed based on the given text prompt; a hierarchical image selection module is introduced to find the most suitable image returned by an image search engine to enhance the T2I model; a self-reflection mechanism is presented to continuously evaluate and refine the generated image to ensure faithful alignment with the text prompt. To evaluate the proposed framework's performance, we collect a dataset named Img-Ref-T2I, where text prompts include three types of uncertain knowledge: (1) known but rare. (2) unknown. (3) ambiguous. Moreover, we carefully craft a complex prompt to guide GPT-4o in making preference evaluation, which has been shown to have an evaluation accuracy similar to that of human preference evaluation. Experimental results demonstrate the effectiveness of our framework, outperforming GPT-4o by about 30% in human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。