arXiv:2601.01618cs.RO2026-01被引 5

用可视草图让机器人边想边画,长程操作更准更可解释。

Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation

  • 引入视觉草图外化空间意图,连接语言与场景几何
  • 循环式看-想-画-动流程提升动态环境适应力
  • 草图可编辑,适合需要人机协作的复杂任务

长程机器人操作在真实场景中日益重要,需在复杂布局中进行空间消歧,并在动态交互中保持时间鲁棒性。现有端到端及分层视觉-语言-动作(VLA)策略常仅依赖文本线索,使规划意图隐含,导致在杂乱或描述不足场景中指代不明确,难以有效分解长程目标,且阻碍因果解释。为此,我们提出视觉草图(Visual Sketch),一种非真实但结构化的视觉中间表示,可渲染点、框、箭头和类型化关系,将机器人视角中的空间意图外化并连接语言与场景几何。基于此,我们构建了Action-Sketcher框架,采用循环式“看-想-画-动”工作流,由自适应令牌门控策略协调推理触发、草图修正与动作执行,支持实时反应与人机交互,同时保持动作预测效率。为实现可扩展训练与评估,我们构建了包含图像、文本、视觉草图标注与动作序列的多样化语料库,并采用多阶段课程学习:融合模态对齐、语言-草图一致性约束,以及增强草图到动作的强化学习模仿学习。在模拟与真实场景的杂乱环境与多物体任务中,实验表明该方法显著提升长程成功率,增强对动态变化的鲁棒性,并通过可编辑草图与分步计划提升可解释性。

原文摘要 · Abstract (English)

Long-horizon robotic manipulation is increasingly important for real-world deployment, requiring spatial disambiguation in complex layouts and temporal resilience under dynamic interaction. However, existing end-to-end and hierarchical Vision-Language-Action (VLA) policies often rely on text-only cues while keeping plan intent latent, which undermines referential grounding in cluttered or underspecified scenes, impedes effective task decomposition of long-horizon goals with close-loop interaction, and limits causal explanation by obscuring the rationale behind action choices. To address these issues, we first introduce Visual Sketch, an implausible visual intermediate that renders points, boxes, arrows, and typed relations in the robot's current views to externalize spatial intent, connect language to scene geometry. Building on Visual Sketch, we present Action-Sketcher, a VLA framework that operates in a cyclic See-Think-Sketch-Act workflow coordinated by adaptive token-gated strategy for reasoning triggers, sketch revision, and action issuance, thereby supporting reactive corrections and human interaction while preserving real-time action prediction. To enable scalable training and evaluation, we curate diverse corpus with interleaved images, text, Visual Sketch supervision, and action sequences, and train Action-Sketcher with a multi-stage curriculum recipe that combines interleaved sequence alignment for modality unification, language-to-sketch consistency for precise linguistic grounding, and imitation learning augmented with sketch-to-action reinforcement for robustness. Extensive experiments on cluttered scenes and multi-object tasks, in simulation and on real-world tasks, show improved long-horizon success, stronger robustness to dynamic scene changes, and enhanced interpretability via editable sketches and step-wise plans. Project website: https://action-sketcher.github.io

机器人操作视觉草图可解释性长程任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。