用强化学习让AI自动拆解复杂图像指令并组合多个模型完成创作。
Image-POSER: Reflective RL for Multi-Expert Image Generation and Editing
- 把图像生成任务看作决策过程,动态拆解长指令并调用不同模型协作。
- 在标准和自定义测试中,生成效果在对齐度、清晰度和美感上均超越现有模型。
- 适合需要复杂图像创作与编辑的设计师或研究人员使用。
文本到图像生成的最新进展已催生出强大的单次生成模型,但尚无单一系统能可靠执行创意工作流中常见的长且复合的指令。我们提出 Image-POSER,一种反射式强化学习框架,(i) 协调一个多样化的预训练文本到图像与图像到图像专家库,(ii) 通过动态任务分解实现对长格式提示的端到端处理,(iii) 利用视觉语言模型批评者提供结构化反馈,在每一步监督对齐。将图像合成与编辑建模为马尔可夫决策过程,我们学习到非平凡的专家流水线,能够自适应地结合各模型优势。实验表明,Image-POSER 在行业标准及自定义基准上均优于基线模型,包括前沿模型,在对齐度、保真度和美学表现方面表现优异,并在人工评估中持续更受青睐。结果表明,强化学习可赋予AI系统自主分解、重排与组合视觉模型的能力,迈向通用视觉助手。
原文摘要 · Abstract (English)
Recent advances in text-to-image generation have produced strong single-shot models, yet no individual system reliably executes the long, compositional prompts typical of creative workflows. We introduce Image-POSER, a reflective reinforcement learning framework that (i) orchestrates a diverse registry of pretrained text-to-image and image-to-image experts, (ii) handles long-form prompts end-to-end through dynamic task decomposition, and (iii) supervises alignment at each step via structured feedback from a vision-language model critic. By casting image synthesis and editing as a Markov Decision Process, we learn non-trivial expert pipelines that adaptively combine strengths across models. Experiments show that Image-POSER outperforms baselines, including frontier models, across industry-standard and custom benchmarks in alignment, fidelity, and aesthetics, and is consistently preferred in human evaluations. These results highlight that reinforcement learning can endow AI systems with the capacity to autonomously decompose, reorder, and combine visual models, moving towards general-purpose visual assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。