提出四层智能生成控制框架,厘清何时视觉生成才算真正‘自主’。
Agentic Visual Generation: From Generative Models to Agentic Control

- 按控制器能掌控的生成环节,划分从输入到经验复用的四层控制能力
- 发现当前系统多在L2-L3层徘徊,缺乏对决策范围的统一标准
- 适合研究生成系统自主性或设计智能创作工具的开发者
视觉生成正从单次调用的生成模型,转向具备规划、选工具、检查中间结果、修正失败和复用经验的智能体控制流程。现有系统通常以LLM或VLM为控制器,视觉生成模型作为执行工具。但缺乏判断系统是否真正具备智能体特性的统一标准。规划深度、工具使用、多角色协作、强化学习常被当作智能体证据,却未必决定控制器实际能做出哪些生成决策。本文依据控制器在生成过程中可直接控制的内容,建立分层框架:L1条件控制(仅准备输入)、L2执行控制(选择并调用生成/编辑/渲染等操作)、L3结果自适应控制(根据中间结果调整后续操作)、L4经验自适应控制(利用已完成任务的经验优化未来决策)。L0固定支持指无部署控制器的生成器、编辑器、评估器、奖励模型、基准测试和固定流水线。该框架体现的是决策范围逐步扩大,而非模型规模、系统复杂度、输出质量、工具数量或训练方法。将此框架应用于图像、视频、编辑、3D、世界建模、幻灯片及用户界面生成,揭示了控制器能力的演进路径与机制分布。
原文摘要 · Abstract (English)
Visual generation is evolving from generative models used through a single invocation into agentic control processes that can plan, select tools, inspect intermediate synthesized outputs, revise failures, and reuse prior experience. In most existing systems, the controller is an LLM or VLM, while visual generation models serve as tools or executors. However, existing work lacks a consistent criterion for determining when a generation system becomes agentic. Planning depth, tool use, multi-role collaboration, and reinforcement learning are often treated as evidence of agenticity, even though none of them necessarily determines which generation decisions the controller can make. We organize the field according to what the controller can directly control in the generation process. At L1 Conditioning Control, the controller prepares the input to a predetermined generator but does not control which visual operation is executed. At L2 Execution Control, it selects and invokes actual generation, editing, rendering, or other content-modifying operations. At L3 Outcome-Adaptive Control, it observes an intermediate outcome and uses that observation to change a subsequent operation within the current task. At L4 Experience-Adaptive Control, it retains experience from completed tasks and uses that experience to change decisions on future tasks. L0 Fixed Support separately denotes generators, editors, evaluators, reward models, benchmarks, and fixed pipelines without a deployed controller that makes generation-level decisions. These levels describe a progressively broader decision-making scope rather than model size, system complexity, output quality, tool or role count, or training method. Applying this framework across image, video, editing, 3D, world, slide, and user-interface generation reveals how controller capabilities have evolved and how their mechanisms are distributed across levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。