用大模型自动生成ComfyUI创意工作流,省去手动配置的复杂步骤。
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
- 基于4000个真实工作流构建思维链数据,训练出首个专用推理模型。
- 7B参数模型格式正确率达97%,性能超越GPT-4o等闭源模型。
- 擅长生成包含多种节点的复杂工作流,适合艺术创作自动化场景。
AI生成内容已从单一模型转向模块化工作流,尤其在ComfyUI等平台实现创意流程的定制化。但设计高效工作流需深厚经验,学习门槛高。为此,我们提出ComfyUI-R1,首个用于自动工作流生成的大规模推理模型。基于自建的4000个工作流数据集,构建包含节点选择、流程规划与代码级表示的长链式思维(CoT)数据。采用两阶段训练框架:(1) 通过思维链微调实现冷启动,使模型适配ComfyUI领域;(2) 基于细粒度规则-指标混合奖励机制进行强化学习,确保格式有效性、结构完整性和节点级一致性。实验表明,7B参数模型达到97%的格式有效率,各项指标显著优于使用GPT-4o、Claude系列等先进闭源模型的现有方法。进一步分析揭示推理过程的关键作用及将工作流转化为代码的优势。定性对比显示,该模型在整合多样节点生成复杂工作流方面表现突出,凸显长链思维推理在AI艺术创作中的潜力。
原文摘要 · Abstract (English)
AI-generated content has evolved from monolithic models to modular workflows, particularly on platforms like ComfyUI, enabling customization in creative pipelines. However, crafting effective workflows requires great expertise to orchestrate numerous specialized components, presenting a steep learning curve for users. To address this challenge, we introduce ComfyUI-R1, the first large reasoning model for automated workflow generation. Starting with our curated dataset of 4K workflows, we construct long chain-of-thought (CoT) reasoning data, including node selection, workflow planning, and code-level workflow representation. ComfyUI-R1 is trained through a two-stage framework: (1) CoT fine-tuning for cold start, adapting models to the ComfyUI domain; (2) reinforcement learning for incentivizing reasoning capability, guided by a fine-grained rule-metric hybrid reward, ensuring format validity, structural integrity, and node-level fidelity. Experiments show that our 7B-parameter model achieves a 97\% format validity rate, along with high pass rate, node-level and graph-level F1 scores, significantly surpassing prior state-of-the-art methods that employ leading closed-source models such as GPT-4o and Claude series. Further analysis highlights the critical role of the reasoning process and the advantage of transforming workflows into code. Qualitative comparison reveals our strength in synthesizing intricate workflows with diverse nodes, underscoring the potential of long CoT reasoning in AI art creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。