arXiv:2509.19571cs.ROcs.CV2025-09被引 2

用场景智能体统一空间、语义与物体功能,让机器人零样本理解复杂指令

Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action

  • 构建可查询的场景表征,结合语义、空间与物体功能进行推理
  • 在桌面上实现零样本语言指令执行,支持房间级任务规划
  • 适合需要灵活应对新场景的机器人应用,如家庭服务与仓储

执行开放式自然语言指令是机器人领域的一个核心挑战。尽管模仿学习和视觉-语言-动作模型(VLAs)取得了进展,但在面对复杂指令或新环境时仍表现不佳。本文提出一种名为“智能体化场景策略”(Agentic Scene Policies, ASP)的框架,通过现代场景表示的语义、空间与功能查询能力,构建一个机器人与环境之间的可查询接口,以指导后续运动规划。ASP 能以零样本方式执行开放词汇查询,尤其在处理复杂技能时显式推理物体功能。通过大量实验,我们对比了 ASP 与 VLAs 在桌面操作任务上的表现,验证了 ASP 可通过功能引导导航完成房间级任务,并支持大规模场景表示。

原文摘要 · Abstract (English)

Executing open-ended natural language queries is a core problem in robotics. While recent advances in imitation learning and vision-language-actions models (VLAs) have enabled promising end-to-end policies, these models struggle when faced with complex instructions and new scenes. An alternative is to design an explicit scene representation as a queryable interface between the robot and the world, using query results to guide downstream motion planning. In this work, we present Agentic Scene Policies (ASP), an agentic framework that leverages the advanced semantic, spatial, and affordance-based querying capabilities of modern scene representations to implement a capable language-conditioned robot policy. ASP can execute open-vocabulary queries in a zero-shot manner by explicitly reasoning about object affordances in the case of more complex skills. Through extensive experiments, we compare ASP with VLAs on tabletop manipulation problems and showcase how ASP can tackle room-level queries through affordance-guided navigation, and a scaled-up scene representation. (Project page: https://montrealrobotics.ca/agentic-scene-policies.github.io/)

机器人语言指令场景理解零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。