让机器人零样本完成复杂桌面上的物品操作任务。
ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning

- 通过工作流推理与视觉掩码接口,实现语义到动作的精准映射。
- 在随机场景中达成50%方程构建成功率、70%约束检索成功率。
- 无需重训练即可适应新指令和物理失败,适合复杂交互场景。
开放式的桌面操作要求智能体不仅理解自然语言,还需适应动态环境与执行失败。本文提出ACE(Agentic Control for Embodied Manipulation),一种基于零样本工作流推理的桌面拾放框架。不同于直接低层动作映射,ACE结合代理式工作流推理与两种机器人可执行技能:视觉定位接口与可复用的拾放原语。为连接语义推理与物理控制,当前子目标通过掩码媒介视觉-动作接口进行定位。该统一掩码指定目标物体与放置位置,可随时间追踪、供人验证,并传递给无任务依赖的下游策略执行。关键在于,ACE在多时标记忆支持下形成闭环。动作执行后,系统自动验证子目标是否成功,依据结果推进、重试、修复或重规划。这使系统能在线适应用户修正、场景变化及物理故障。我们在逻辑复杂的长时序任务上评估了ACE,包括使用数字积木的零样本多步方程构建与基于约束的对象检索。结果显示,尽管传统端到端基线难以完成这些高难度任务,ACE在方程构建任务中达到50%成功率,在约束检索任务中达70%成功率。这表明显式工作流推理与掩码媒介控制为可适应的机器人操作提供了稳健、实用的路径。
原文摘要 · Abstract (English)
Open-ended tabletop manipulation requires agents to not only understand natural language but also adapt to dynamic environments and execution failures. We present ACE (Agentic Control for Embodied Manipulation), a zero-shot workflow reasoning framework for tabletop pick-and-place from natural language. Rather than relying on direct low-level action mapping, ACE combines agentic workflow reasoning with two robot-facing executable skills: a visual grounding interface and a reusable pick-and-place primitive. To bridge semantic reasoning and physical control, the active sub-goal is grounded into a mask-mediated vision-action interface. This unified mask specifies the target object and destination, is tracked over time, exposed for human verification, and ultimately passed to a task-agnostic downstream policy for execution. Crucially, ACE operates in a closed loop supported by a multi-timescale memory. After an action is executed, the system automatically verifies whether the intended sub-goal succeeded, using the outcome to advance, retry, repair, or replan. This enables online adaptation to user corrections, scene changes, and physical failures. We evaluate ACE on logically complex, long-horizon tasks, including zero-shot multi-step equation formation with number cubes and constraint-based object retrieval. ACE demonstrates task-level zero-shot generalization on novel semantic constraints and randomized tabletop scenes without task-specific retraining. Specifically, while standard end-to-end baselines struggle to complete these logically demanding tasks, ACE achieves a 50% success rate in equation formation and a 70% success rate in constraint retrieval. This contrast demonstrates that explicit workflow reasoning and mask-mediated control offer a robust, practical route toward adaptable robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。