arXiv:2503.12579cs.ROcs.LG2025-03中稿 · RLDM 2025被引 1

让机器人根据用户目标聚焦学习,避免无意义探索。

Focusing Robot Open-Ended Reinforcement Learning Through Users' Purposes

  • 用用户目的引导自主学习,聚焦相关物体
  • 通过语言理解与场景分析,自动识别关键物体
  • 适合需要个性化交互的智能服务机器人

开放式学习(OEL)自主机器人可通过与环境直接互动获取新技能,依赖内在动机和自生成目标驱动学习。这类机器人在非结构化环境中能自主利用知识完成对用户有益的任务,应对设计时未预见的挑战。然而,其开放性可能导致学习无关信息。本文提出‘目的导向开放式学习’(POEL),基于先前工作提出的‘目的’概念——即用户希望机器人达成的目标。核心思想是:目的可引导OEL聚焦于自生成任务类别,这些任务虽在自主学习中未知,但涉及与目的相关的物体。该方法在新型机器人架构中实现,支持语音转文本接收人类目的,通过场景分析识别物体,并利用大语言模型判断物体相关性。相关物体被用于引导探索方向并生成奖励,鼓励与它们交互。在摄像头-机械臂-夹爪机器人模拟场景中验证,结果首次证明目的聚焦的OEL优于现有先进OEL方法,使机器人能在非结构化环境中有效学习,同时保持与用户需求一致的知识获取。

原文摘要 · Abstract (English)

Open-Ended Learning (OEL) autonomous robots can acquire new skills and knowledge through direct interaction with their environment, relying on mechanisms such as intrinsic motivations and self-generated goals to guide learning processes. OEL robots are highly relevant for applications as they can autonomously leverage acquired knowledge to perform tasks beneficial to human users in unstructured environments, addressing challenges unforeseen at design time. However, OEL robots face a significant limitation: their openness may lead them to waste time learning information that is irrelevant to tasks desired by specific users. Here, we propose a solution called `Purpose-Directed Open-Ended Learning' (POEL), based on the novel concept of `purpose' introduced in previous work. A purpose specifies what users want the robot to achieve. The key insight of this work is that purpose can focus OEL on learning self-generated classes of tasks that, while unknown during autonomous learning (as typical in OEL), involve objects relevant to the purpose. This concept is operationalised in a novel robot architecture capable of receiving a human purpose through speech-to-text, analysing the scene to identify objects, and using a Large Language Model to reason about which objects are purpose-relevant. These objects are then used to bias OEL exploration towards their spatial proximity and to self-generate rewards that favour interactions with them. The solution is tested in a simulated scenario where a camera-arm-gripper robot interacts freely with purpose-related and distractor objects. For the first time, the results demonstrate the potential advantages of purpose-focused OEL over state-of-the-art OEL methods, enabling robots to handle unstructured environments while steering their learning toward knowledge acquisition relevant to users.

机器人学习开放式学习用户意图大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。