arXiv:2512.19243cs.CV2025-12被引 2

让AI理解复杂设计指令,自动修正生成图像中的细节错误。

VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis

  • 用视觉语言模型解析长篇设计指令,分步执行编辑
  • 在2000个任务中实现72%以上目标达成率,优于现有模型
  • 适合需要精准图文生成的设计师和工业级应用

当前生成模型虽能产出逼真图像,但在处理专业设计师提出的长周期、多目标指令时仍表现不佳。为此,我们构建了包含2000项任务(1000项文本到图像、1000项图像到图像)的Long Goal Bench(LGBench),其平均指令包含18至22个紧密耦合的目标,涵盖整体布局、局部物体位置、排版与标识一致性。实验发现,即使顶尖模型也仅满足不足72%的目标,且常遗漏局部修改。为解决该问题,我们提出VisionDirector——一种无需训练的视觉语言监督框架:(i) 从长指令中提取结构化目标;(ii) 动态决策一次性生成或分步编辑;(iii) 每次编辑后进行微网格采样、语义验证及回滚;(iv) 记录目标级奖励。进一步通过群组相对策略优化微调规划器,使编辑轨迹缩短至3.1步(原4.2步),提升对齐度。VisionDirector在GenEval上提升7%总体得分,在ImgEdit上提升0.07绝对分数,且在排版、多对象场景与姿态编辑方面均实现一致的定性改进。

原文摘要 · Abstract (English)

Generative models can now produce photorealistic imagery, yet they still struggle with the long, multi-goal prompts that professional designers issue. To expose this gap and better evaluate models' performance in real-world settings, we introduce Long Goal Bench (LGBench), a 2,000-task suite (1,000 T2I and 1,000 I2I) whose average instruction contains 18 to 22 tightly coupled goals spanning global layout, local object placement, typography, and logo fidelity. We find that even state-of-the-art models satisfy fewer than 72 percent of the goals and routinely miss localized edits, confirming the brittleness of current pipelines. To address this, we present VisionDirector, a training-free vision-language supervisor that (i) extracts structured goals from long instructions, (ii) dynamically decides between one-shot generation and staged edits, (iii) runs micro-grid sampling with semantic verification and rollback after every edit, and (iv) logs goal-level rewards. We further fine-tune the planner with Group Relative Policy Optimization, yielding shorter edit trajectories (3.1 versus 4.2 steps) and stronger alignment. VisionDirector achieves new state of the art on GenEval (plus 7 percent overall) and ImgEdit (plus 0.07 absolute) while producing consistent qualitative improvements on typography, multi-object scenes, and pose editing.

图像生成视觉语言闭环优化设计辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。