arXiv:2609.05369cs.ROcs.CV2026-09

让AI机器人更可靠地完成复杂长流程操作任务。

Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

论文配图:Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation
图 1 · 摘自论文原文
  • 结合符号逻辑与深度学习,用任务图和记忆模块管理步骤依赖。
  • 在清理工作区和手术器械处理任务中,成功率提升至85%以上。
  • 适合需要精确步骤控制的机器人操作场景研究者使用。

视觉-语言-动作(VLA)模型可执行短时操作技能,但在需要持续任务状态、依赖感知推理、条件决策和可靠定位的长周期流程中仍显脆弱。本文提出一种神经符号框架,融合学习到的VLA控制与显式任务图及多模态过程记忆。任务图编码动作依赖、有效转移和分支条件,记忆则维护当前步骤、已完成动作、文本上下文及任务相关的视觉证据。这些结构共同指导物体选择、目标定位、子目标分发与预期状态转移验证。人类示范通过注视或显著性线索提供额外时空引导,研究中通过机器人视角遥操作视频直接标注伪注视以隔离其对策略学习的影响,该引导用于VLA微调与推理。我们在工作区清理和手术器械处理两个长周期操作领域进行评估,涵盖正确物体与目标选择、子任务完成、任务进展、步骤顺序一致性、完整任务成功及程序性错误等指标。结果表明,结构化符号推理与示范衍生的视觉引导是实现可靠长周期VLA操作的互补机制。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investigate a neuro-symbolic framework that combines learned VLA control with explicit task graphs and multimodal procedural memory. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory maintains the active step, completed actions, textual context, and task-relevant visual evidence. Together, these structures guide object selection, destination grounding, subgoal dispatch, and verification of expected state transitions. Human demonstrations provide additional spatial and temporal guidance through gaze or saliency cues. To isolate their effect on policy learning, our initial study bypasses cross-view gaze transfer and directly annotates pseudo-gaze in robot-view teleoperation videos. The resulting guidance is used during VLA fine-tuning and inference. We study two long-horizon manipulation domains, workspace clearing and surgical-instrument handling, which require ordered execution, visually grounded decisions, and conditional branching. We evaluate correct-object and destination selection, subtask completion, task progress, step-order consistency, complete-task success, and procedural or execution mistakes. This work positions structured symbolic reasoning and demonstration-derived visual guidance as complementary mechanisms for reliable long-horizon VLA manipulation.

机器人操作符号推理长程任务多模态记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。