提出一种简单有效的技能提取终止条件,自动识别关键决策点。
NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations
- 用状态-动作新颖性模块分析经验数据,自动判断何时终止技能。
- 在复杂长任务中表现优于基线,环境变化时仍保持稳定性能。
- 适合需要自适应技能拆分的强化学习场景,如机器人操作。
智能代理能够基于不同粒度和持续时间做出决策。近年来的技能学习进展使代理能通过引导选择合适技能来解决复杂、长周期任务。然而,固定长度的技能容易跳过有价值的决策点,限制了进一步探索和更快策略学习的潜力。本文提出一种简单有效的终止条件,通过状态-动作新颖性模块利用代理的经验数据识别决策点。所提方法名为基于新颖性的决策点识别(NBDI),在复杂长任务中优于先前基线,在下游任务环境配置存在显著变化时仍保持有效性,凸显了决策点识别在技能学习中的重要性。
原文摘要 · Abstract (English)
Intelligent agents are able to make decisions based on different levels of granularity and duration. Recent advances in skill learning enabled the agent to solve complex, long-horizon tasks by effectively guiding the agent in choosing appropriate skills. However, the practice of using fixed-length skills can easily result in skipping valuable decision points, which ultimately limits the potential for further exploration and faster policy learning. In this work, we propose to learn a simple and effective termination condition that identifies decision points through a state-action novelty module that leverages agent experience data. Our approach, Novelty-based Decision Point Identification (NBDI), outperforms previous baselines in complex, long-horizon tasks, and remains effective even in the presence of significant variations in the environment configurations of downstream tasks, highlighting the importance of decision point identification in skill learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。