arXiv:2603.02681cs.CV2026-03被引 1

提出可自主创作的视觉生成智能体,统一理解、思考、规划与创作能力。

VisionCreator: A Native Visual-Generation Agentic Model with Understanding, Thinking, Planning and Creation

  • 构建具思维结构的视觉创作数据集,支持端到端训练
  • 在模拟环境中通过渐进式训练获得复杂创作能力
  • 性能超越更大闭源模型,适合创意设计研究者使用

视觉内容创作需要对设计规范和创作流程有深入理解,通用模型难以胜任;而基于工作流的智能体又缺乏自主创作所需的专门知识。为此,我们提出 VisionCreator,一种原生视觉生成智能体,将理解、思考、规划与创作(UTPC)能力统一于端到端可学习框架中。本工作提出四项关键贡献:(i) 构建包含4000条高质量创作轨迹的 VisGenData-4k 数据集,采用基于元认知的 VisionAgent 生成具有明确 UTPC 结构的轨迹;(ii) VisionCreator 智能体通过渐进式专业化训练(PST)和虚拟强化学习(VRL),在高保真模拟环境中优化,实现复杂创作任务下稳定高效的 UTPC 能力获取;(iii) 设计 VisGenBench 基准,涵盖1200个跨场景测试样本,用于标准化评估多步视觉生成能力;(iv) 实验表明,VisionCreator-8B/32B 模型在多个维度上表现优于更大规模的闭源模型。本工作为视觉生成智能体系统研究奠定基础。

原文摘要 · Abstract (English)

Visual content creation tasks demand a nuanced understanding of design conventions and creative workflows-capabilities challenging for general models, while workflow-based agents lack specialized knowledge for autonomous creative planning. To overcome these challenges, we propose VisionCreator, a native visual-generation agentic model that unifies Understanding, Thinking, Planning, and Creation (UTPC) capabilities within an end-to-end learnable framework. Our work introduces four key contributions: (i) VisGenData-4k and its construction methodology using metacognition-based VisionAgent to generate high-quality creation trajectories with explicit UTPC structures; (ii) The VisionCreator agentic model, optimized through Progressive Specialization Training (PST) and Virtual Reinforcement Learning (VRL) within a high-fidelity simulated environment, enabling stable and efficient acquisition of UTPC capabilities for complex creation tasks; (iii) VisGenBench, a comprehensive benchmark featuring 1.2k test samples across diverse scenarios for standardized evaluation of multi-step visual creation capabilities; (iv) Remarkably, our VisionCreator-8B/32B models demonstrate superior performance over larger closed-source models across multiple evaluation dimensions. Overall, this work provides a foundation for future research in visual-generation agentic systems.

视觉生成智能体创作系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。