arXiv:2504.14805cs.LGcs.AI2025-04ICLR被引 3

动态调整技能长度,让强化学习更灵活地识别复杂行为。

Dynamic Contrastive Skill Learning with State-Transition Based Skill Clustering and Dynamic Length Adjustment

  • 基于状态转移构建技能表示,捕捉行为语义上下文。
  • 通过对比学习自动识别相似行为,提升技能泛化能力。
  • 适合处理复杂或噪声数据,特别适用于长时序任务。

强化学习在多个领域已取得显著进展,但将其扩展到需要复杂决策的长时序任务仍具挑战性。技能学习通过将动作抽象为高层行为来应对这一问题。然而,现有方法常无法将语义相似的行为识别为同一技能,且使用固定技能长度,限制了灵活性和泛化能力。为此,我们提出动态对比技能学习(DCSL),一种重新定义技能表示与学习的新框架。DCSL引入三个核心思想:基于状态转移的技能表示、技能相似性函数学习以及动态技能长度调整。通过聚焦状态转移并利用对比学习,DCSL有效捕捉行为的语义上下文,并根据行为的时间跨度自适应调整技能长度。该方法在复杂或噪声数据中表现出更强的灵活性与适应性,在任务完成率与效率方面优于现有方法。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has made significant progress in various domains, but scaling it to long-horizon tasks with complex decision-making remains challenging. Skill learning attempts to address this by abstracting actions into higher-level behaviors. However, current approaches often fail to recognize semantically similar behaviors as the same skill and use fixed skill lengths, limiting flexibility and generalization. To address this, we propose Dynamic Contrastive Skill Learning (DCSL), a novel framework that redefines skill representation and learning. DCSL introduces three key ideas: state-transition based skill representation, skill similarity function learning, and dynamic skill length adjustment. By focusing on state transitions and leveraging contrastive learning, DCSL effectively captures the semantic context of behaviors and adapts skill lengths to match the appropriate temporal extent of behaviors. Our approach enables more flexible and adaptive skill extraction, particularly in complex or noisy datasets, and demonstrates competitive performance compared to existing methods in task completion and efficiency.

强化学习技能学习动态长度对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。