arXiv:2505.02232cs.ROcs.AI2025-05ICRA被引 2

让机器人听懂指令精准抓取杂乱中的物体

Prompt-responsive Object Retrieval with Memory-augmented Student-Teacher Learning

  • 用记忆增强的师生学习框架连接指令与操作
  • 在杂乱场景中成功实现指令响应的抓取任务
  • 适合做具身智能、人机交互方向的研究者参考

构建对输入指令敏感的模型是机器学习的一次范式转变。该方法在机器人领域具有重要意义,例如在杂乱环境中进行目标操作。本文提出一种新方法,将可提示的基础模型与强化学习结合,使机器人能够以指令响应方式完成灵巧操作任务。现有方法难以将高层指令与精细控制关联。为此,我们设计了记忆增强的师生学习框架,采用Segment-Anything 2(SAM 2)作为感知主干,从用户指令中推断目标物体。尽管检测不完美,但其时序序列为记忆增强模型提供了丰富的隐状态估计信息。实验表明,该方法成功学习到指令响应策略,在从杂乱场景中抓取物体的任务上表现良好。视频与代码见 https://memory-student-teacher.github.io

原文摘要 · Abstract (English)

Building models responsive to input prompts represents a transformative shift in machine learning. This paradigm holds significant potential for robotics problems, such as targeted manipulation amidst clutter. In this work, we present a novel approach to combine promptable foundation models with reinforcement learning (RL), enabling robots to perform dexterous manipulation tasks in a prompt-responsive manner. Existing methods struggle to link high-level commands with fine-grained dexterous control. We address this gap with a memory-augmented student-teacher learning framework. We use the Segment-Anything 2 (SAM 2) model as a perception backbone to infer an object of interest from user prompts. While detections are imperfect, their temporal sequence provides rich information for implicit state estimation by memory-augmented models. Our approach successfully learns prompt-responsive policies, demonstrated in picking objects from cluttered scenes. Videos and code are available at https://memory-student-teacher.github.io

机器人操作指令响应视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。