无需更新模型,仅用少量数据即可快速适配新环境的策略框架。
In-Context Policy Adaptation via Cross-Domain Skill Diffusion
- 通过跨域技能扩散联合学习通用技能原型与领域适配器。
- 在有限目标数据下,对不同环境动态、机器人形态和任务时长均表现优异。
- 适合需要快速部署且无法更新模型的机器人或自动驾驶场景。
本文提出一种面向长时程多任务环境的上下文策略自适应(ICPAD)框架,探索基于扩散的跨域技能学习方法。该框架在禁止模型更新且仅有少量目标域数据的严苛条件下,实现技能型强化学习策略的快速适应。通过离线数据联合学习通用技能原型与领域感知的技能适配器,利用跨域一致性扩散过程确保技能可迁移性。通用技能原型作为长时程策略的通用行为表征,充当跨域沟通桥梁。为提升上下文适应性能,设计动态领域提示机制,引导扩散式适配器更优对齐目标域。在Metaworld机器人操作和CARLA自动驾驶场景中的实验表明,本框架在多种跨域配置(包括环境动力学差异、代理体态差异与任务时长差异)下,均在有限目标域数据条件下实现了优越的策略适应性能。
原文摘要 · Abstract (English)
In this work, we present an in-context policy adaptation (ICPAD) framework designed for long-horizon multi-task environments, exploring diffusion-based skill learning techniques in cross-domain settings. The framework enables rapid adaptation of skill-based reinforcement learning policies to diverse target domains, especially under stringent constraints on no model updates and only limited target domain data. Specifically, the framework employs a cross-domain skill diffusion scheme, where domain-agnostic prototype skills and a domain-grounded skill adapter are learned jointly and effectively from an offline dataset through cross-domain consistent diffusion processes. The prototype skills act as primitives for common behavior representations of long-horizon policies, serving as a lingua franca to bridge different domains. Furthermore, to enhance the in-context adaptation performance, we develop a dynamic domain prompting scheme that guides the diffusion-based skill adapter toward better alignment with the target domain. Through experiments with robotic manipulation in Metaworld and autonomous driving in CARLA, we show that our $\oursol$ framework achieves superior policy adaptation performance under limited target domain data conditions for various cross-domain configurations including differences in environment dynamics, agent embodiment, and task horizon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。