arXiv:2603.08812cs.CV2026-03被引 1

提出可自我反思的视觉生成智能体,提升多图创作准确性

VisionCreator-R1: A Reflection-Enhanced Native Visual-Generation Agentic Model

  • 引入显式反思机制,让生成过程能中途纠错
  • 在单图和多图任务上均超越Gemini 2.5 Pro表现
  • 适合需要高质量图像创作的AI应用开发者

视觉内容生成已从单图发展到多图工作流,但现有智能体仍以计划驱动为主,缺乏系统性反思机制来修正生成过程中的视觉错误。为此,我们提出VisionCreator-R1,一种具备显式反思能力的原生视觉生成智能体,并设计了反射-计划协同优化(RPCO)训练方法。通过大量实验与轨迹级分析,我们发现强化学习中存在反思-计划优化不对称:计划可通过计划奖励可靠优化,而反思学习受噪声信用分配阻碍。基于此洞察,RPCO首先在自构建的VCR-SFT数据集上训练,分别使用反思强的单图轨迹和计划强的多图轨迹;随后在VCR-RL数据集上进行强化学习协同优化。最终得到统一的VisionCreator-R1智能体,在现有基准及涵盖单图与多图任务的VCR-bench上持续优于Gemini 2.5 Pro。

原文摘要 · Abstract (English)

Visual content generation has advanced from single-image to multi-image workflows, yet existing agents remain largely plan-driven and lack systematic reflection mechanisms to correct mid-trajectory visual errors. To address this limitation, we propose VisionCreator-R1, a native visual generation agent with explicit reflection, together with a Reflection-Plan Co-Optimization (RPCO) training methodology. Through extensive experiments and trajectory-level analysis, we uncover reflection-plan optimization asymmetry in reinforcement learning (RL): planning can be reliably optimized via plan rewards, while reflection learning is hindered by noisy credit assignment. Guided by this insight, our RPCO first trains on the self-constructed VCR-SFT dataset with reflection-strong single-image trajectories and planning-strong multi-image trajectories, then co-optimization on VCR-RL dataset via RL. This yields our unified VisionCreator-R1 agent, which consistently outperforms Gemini2.5Pro on existing benchmarks and our VCR-bench covering single-image and multi-image tasks.

视觉生成智能体反思机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。