arXiv:2603.19500cs.AIcs.CV2026-03被引 1

让智能体分步绘制草图,支持局部修改与可控生成。

Teaching an Agent to Sketch One Part at a Time

  • 用多轮过程奖励强化学习训练语言模型代理,逐部分生成草图。
  • 引入新数据集ControlSketch-Part,实现草图的语义部件标注。
  • 支持交互式编辑,生成结果可解释、可控且局部可调。

我们提出一种分步生成矢量草图的方法。通过在监督微调后采用新型多轮过程奖励强化学习训练基于多模态语言模型的智能体,实现逐部分绘制。该方法依托新构建的数据集ControlSketch-Part,该数据集通过通用自动标注流程,将矢量草图分割为语义部件,并通过多阶段结构化标注流程为路径分配部件标签。实验表明,结合结构化的部件级数据并提供生成过程中的视觉反馈,可实现可解释、可控且支持局部编辑的文本到矢量草图生成。

原文摘要 · Abstract (English)

We develop a method for producing vector sketches one part at a time. To do this, we train a multi-modal language model-based agent using a novel multi-turn process-reward reinforcement learning following supervised fine-tuning. Our approach is enabled by a new dataset we call ControlSketch-Part, containing rich part-level annotations for sketches, obtained using a novel, generic automatic annotation pipeline that segments vector sketches into semantic parts and assigns paths to parts with a structured multi-stage labeling process. Our results indicate that incorporating structured part-level data and providing agent with the visual feedback through the process enables interpretable, controllable, and locally editable text-to-vector sketch generation.

草图生成分步绘制可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。