让机器人主动选学什么、向谁学,提升多任务学习效率。
The intrinsic motivation of reinforcement and imitation learning for sequential tasks
- 基于经验进展设计内在动机机制,自动选择学习策略
- 比被动接收演示快,减少所需示范次数,抗导师质量波动
- 适合需要高效模仿与自主学习的机器人系统研究
本研究面向发展认知机器人领域,提出一种融合强化学习与模仿学习的新框架,通过内在动机引导学习者在多个任务(包括序列任务)中主动选择学习路径。核心贡献是建立基于经验进展的统一内在动机模型,使学习者可自动决定:学哪个任务、是否自主探索或模仿、使用底层动作还是任务分解、选择哪位导师。该机制不仅被动接收演示,更主动请求最优导师的示范,显著提升对导师质量的鲁棒性,并加速学习过程。我们构建了社会引导的内在动机框架,结合机器学习算法,在利用人类示范泛化能力的基础上,支持被动与主动请求两种模式。尤其针对组合子任务,提出子任务构造成分表示,需结合对日常活动观测的表示进行优化。展望与导师的语言式交互,探索了连续感知运动空间与任务的符号表征生成。在强化学习框架下,提出与导师互动的奖励函数,实现多任务学习中的自动课程学习。
原文摘要 · Abstract (English)
This work in the field of developmental cognitive robotics aims to devise a new domain bridging between reinforcement learning and imitation learning, with a model of the intrinsic motivation for learning agents to learn with guidance from tutors multiple tasks, including sequential tasks. The main contribution has been to propose a common formulation of intrinsic motivation based on empirical progress for a learning agent to choose automatically its learning curriculum by actively choosing its learning strategy for simple or sequential tasks: which task to learn, between autonomous exploration or imitation learning, between low-level actions or task decomposition, between several tutors. The originality is to design a learner that benefits not only passively from data provided by tutors, but to actively choose when to request tutoring and what and whom to ask. The learner is thus more robust to the quality of the tutoring and learns faster with fewer demonstrations. We developed the framework of socially guided intrinsic motivation with machine learning algorithms to learn multiple tasks by taking advantage of the generalisability properties of human demonstrations in a passive manner or in an active manner through requests of demonstrations from the best tutor for simple and composing subtasks. The latter relies on a representation of subtask composition proposed for a construction process, which should be refined by representations used for observational processes of analysing human movements and activities of daily living. With the outlook of a language-like communication with the tutor, we investigated the emergence of a symbolic representation of the continuous sensorimotor space and of tasks using intrinsic motivation. We proposed within the reinforcement learning framework, a reward function for interacting with tutors for automatic curriculum learning in multi-task learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。