arXiv:2507.06628cs.LGcs.AI2025-07ICML被引 3

通过目标导向技能抽象,提升离线多任务强化学习的知识迁移能力。

Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning

  • 基于目标提取可复用技能,构建离散技能库
  • 解决通用与专用技能不平衡问题,提升任务表现
  • 适合需要高效知识共享的机器人控制场景

离线多任务强化学习旨在仅使用预先收集的任务混合数据集,不依赖环境在线交互,学习一个能解决多个任务的统一策略。然而,其在跨任务有效知识共享方面面临挑战。受人类学习中高效知识抽象的启发,我们提出目标导向技能抽象(GO-Skill),通过目标导向的技能提取过程发现可复用技能,并利用向量量化构建离散技能库。为缓解广泛适用技能与任务特定技能之间的类别不平衡问题,引入技能增强阶段优化提取结果。进一步采用分层策略学习整合这些技能,构建高层策略,动态调度离散技能完成具体任务。在MetaWorld基准上的多样化机器人操作任务上进行的大量实验表明,GO-Skill在有效性与泛化性方面均表现出色。

原文摘要 · Abstract (English)

Offline multi-task reinforcement learning aims to learn a unified policy capable of solving multiple tasks using only pre-collected task-mixed datasets, without requiring any online interaction with the environment. However, it faces significant challenges in effectively sharing knowledge across tasks. Inspired by the efficient knowledge abstraction observed in human learning, we propose Goal-Oriented Skill Abstraction (GO-Skill), a novel approach designed to extract and utilize reusable skills to enhance knowledge transfer and task performance. Our approach uncovers reusable skills through a goal-oriented skill extraction process and leverages vector quantization to construct a discrete skill library. To mitigate class imbalances between broadly applicable and task-specific skills, we introduce a skill enhancement phase to refine the extracted skills. Furthermore, we integrate these skills using hierarchical policy learning, enabling the construction of a high-level policy that dynamically orchestrates discrete skills to accomplish specific tasks. Extensive experiments on diverse robotic manipulation tasks within the MetaWorld benchmark demonstrate the effectiveness and versatility of GO-Skill.

强化学习技能抽象离线学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。