arXiv:2607.05465cs.CVcs.AI2026-07

让AI像画家一样用多种工具协作创作复杂图像。

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

论文配图:CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
图 1 · 摘自论文原文
  • 通过多轮交互学习调度不同视觉工具完成创作任务
  • 在14万条轨迹数据上训练,实现高精度图像生成与编辑
  • 适合需要复杂图像创作的设计师和开发者使用

复杂图像创作与编辑往往需要超越单一生成或编辑模型的能力。用户需求可能涉及图像合成、目标定位、区域分割、内容编辑、中间素材拼接、文本识别及最终效果增强等多个步骤。这类任务使多模态智能体从感知增强推理转向以操作为核心的视觉创作,要求工具主动改变视觉状态而非仅观察。然而现有模型多聚焦于感知、搜索或特定领域编辑,缺乏大规模可执行图像创作轨迹的监督数据。本文提出CanvasCraft——一个大规模多模态工具使用数据集,包含14万条完整标注的可执行轨迹和1万条强化学习任务规范,并构建了 extbf{CanvasAgent}:一种通过多轮交互学习协调异构视觉工具的工具增强型多模态智能体。CanvasAgent首先通过SFT学习可执行的推理-动作轨迹,再利用结合结果与过程信号的混合奖励进行GRPO优化。推理过程中,该智能体可检查中间结果、追踪视觉资产,并根据不断变化的视觉状态动态调整工具选择。实验评估了最终图像质量与轨迹行为表现,验证了CanvasAgent及所提数据集在复杂多工具图像创作流程中的有效性。

原文摘要 · Abstract (English)

Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, localizing objects, segmenting regions, editing selected content, compositing intermediate assets, reading text, and enhancing the final result. Such tasks shift multimodal agents from perception-augmented reasoning to manipulation-centered visual creation, where tools must actively transform visual states rather than merely inspect them. However, existing multimodal tool-use agents are mostly optimized for perception, search, or domain-specific editing, and lack large-scale supervision for executable image-creation trajectories. In this paper, we introduce CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and \textbf{CanvasAgent}, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction. CanvasCraft contains 140K fully annotated executable trajectories and 10K RL task specifications. CanvasAgent is first trained with SFT to learn executable reasoning-action trajectories, and is then optimized with GRPO using a hybrid reward that combines outcome- and process-level signals. During rollout, CanvasAgent inspects intermediate results, tracks visual assets, and adapts tool decisions to the evolving visual state. Experiments evaluate both final image quality and trajectory behavior, demonstrating the effectiveness of CanvasAgent and the proposed dataset for complex multi-tool image creation workflows.

图像生成工具调度多模态智能创作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。