根据提示自动生成适配的工作流,提升文生图质量。
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
- 用大模型根据提示选择或生成专用图像生成工作流。
- 相比单一模型和固定流程,图像质量显著提升。
- 适合需要高质量生成的创作者与研究者使用。
文本到图像生成的实际应用已从简单的单体模型演变为结合多个专用组件的复杂工作流。尽管基于工作流的方法能提升图像质量,但设计有效工作流需大量专业知识,因组件数量庞大、相互依赖复杂且受生成提示影响。本文提出新的任务:提示自适应工作流生成,即自动为每个用户提示定制工作流。我们提出两种基于大模型的方法:一种通过用户偏好数据进行微调,另一种无需训练,利用大模型直接选择现有工作流。两种方法均在图像质量上优于单体模型或通用、不依赖提示的工作流。结果表明,提示相关的流程预测为提升文生图质量提供了新路径,补充了该领域的现有研究方向。
原文摘要 · Abstract (English)
The practical use of text-to-image generation has evolved from simple, monolithic models to complex workflows that combine multiple specialized components. While workflow-based approaches can lead to improved image quality, crafting effective workflows requires significant expertise, owing to the large number of available components, their complex inter-dependence, and their dependence on the generation prompt. Here, we introduce the novel task of prompt-adaptive workflow generation, where the goal is to automatically tailor a workflow to each user prompt. We propose two LLM-based approaches to tackle this task: a tuning-based method that learns from user-preference data, and a training-free method that uses the LLM to select existing flows. Both approaches lead to improved image quality when compared to monolithic models or generic, prompt-independent workflows. Our work shows that prompt-dependent flow prediction offers a new pathway to improving text-to-image generation quality, complementing existing research directions in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。