arXiv:2603.02688cs.AIcs.RO2026-03

让机器人像人一样查手册,边查边做复杂组装。

Retrieval-Augmented Robots via Retrieve-Reason-Act

  • 用查文档+理解图示+生成动作的循环提升机器人自主性
  • 在长任务装配中显著优于仅靠推理或回忆旧例子的方法
  • 适合需要查说明书的新任务,尤其缺乏先验经验时

为实现通用功能,机器人需从被动执行者转变为主动信息检索使用者。在无任何示范的零样本场景下,机器人面临关键信息缺失,如组装复杂家具的具体步骤,仅靠常识或内部记忆无法解决。现有方法多检索过往运动轨迹或文本安全规则,无法获取外部非结构化文档中的未见操作知识。本文提出检索增强型机器人(RAR)范式,将任务执行设计为迭代的‘查-想-动’循环:机器人主动从非结构化语料库中检索相关视觉操作手册,通过跨模态对齐将二维图示与三维实物对应,进而生成可执行计划。在一项具有挑战性的长程装配基准上验证,基于检索视觉文档的规划显著优于依赖零样本推理或少量示例检索的基线模型。本工作为信息检索拓展至驱动具身物理行为奠定基础。

原文摘要 · Abstract (English)

To achieve general-purpose utility, we argue that robots must evolve from passive executors into active Information Retrieval users. In strictly zero-shot settings where no prior demonstrations exist, robots face a critical information gap, such as the exact sequence required to assemble a complex furniture kit, that cannot be satisfied by internal parametric knowledge (common sense) or past internal memory. While recent robotic works attempt to use search before action, they primarily focus on retrieving past kinematic trajectories (analogous to searching internal memory) or text-based safety rules (searching for constraints). These approaches fail to address the core information need of active task construction: acquiring unseen procedural knowledge from external, unstructured documentation. In this paper, we define the paradigm as Retrieval-Augmented Robotics (RAR), empowering the robot with the information-seeking capability that bridges the gap between visual documentation and physical actuation. We formulate the task execution as an iterative Retrieve-Reason-Act loop: the robot or embodied agent actively retrieves relevant visual procedural manuals from an unstructured corpus, grounds the abstract 2D diagrams to 3D physical parts via cross-modal alignment, and synthesizes executable plans. We validate this paradigm on a challenging long-horizon assembly benchmark. Our experiments demonstrate that grounding robotic planning in retrieved visual documents significantly outperforms baselines relying on zero-shot reasoning or few-shot example retrieval. This work establishes the basis of RAR, extending the scope of Information Retrieval from answering user queries to driving embodied physical actions.

机器人检索增强视觉理解长程任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。