让统一模型生成时保持视觉一致性,避免人物、物体特征丢失。
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- 用可视化清单规划需保持的视觉特征,引导思考过程。
- 通过自检和迭代修正,提升多参考图像生成的视觉一致性。
- 适合需要高保真视觉输出的多模态生成任务。
链式思维(CoT)虽显著提升了统一模型的生成能力,但现有方法在多模态生成中主要关注文本与提示的一致性,忽视了与视觉参考图像的视觉上下文一致性,导致关键视觉特征(如人物身份、物体属性、风格)无法维持。为此,本文将视觉一致性融入统一模型的推理过程,提出两个核心机制:1)自适应视觉规划,生成结构化视觉检查清单以明确需保持的视觉元素;2)迭代视觉修正,基于检查清单进行自省并逐步优化生成结果。通过监督微调教会模型如何规划、自省与修正,并采用定制化的视觉检查奖励(flow-GRPO)进一步增强一致性。实验表明,该方法在多参考生成任务中优于零样本统一模型及仅使用文本CoT的模型,显著提升了视觉上下文一致性。
原文摘要 · Abstract (English)
Recently, the introduction of Chain-of-Thought (CoT) has largely improved the generation ability of unified models. However, it is observed that the current thinking process during generation mainly focuses on the text consistency with the text prompt, ignoring the \textbf{visual context consistency} with the visual reference images during the multi-modal generation, e.g., multi-reference generation. The lack of such consistency results in the failure in maintaining key visual features (like human ID, object attribute, style). To this end, we integrate the visual context consistency into the reasoning of unified models, explicitly motivating the model to sustain such consistency by 1) Adaptive Visual Planning: generating structured visual check list to figure out the visual element of needed consistency keeping, and 2) Iterative Visual Correction: performing self-reflection with the guidance of check lists and refining the generated result in an iterative manner. To achieve this, we use supervised finetuning to teach the model how to plan the visual checking, conduct self-reflection and self-refinement, and use flow-GRPO to further enhance the visual consistency through a customized visual checking reward. The experiments show that our method outperforms both zero-shot unified models and those with text CoTs in multi-modal generation, demonstrating higher visual context consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。