arXiv:2510.04901cs.LGcs.AI2025-10

让智能体精准控制特定状态变量,避免副作用。

Focused Skill Discovery: Learning to Control Specific State Variables while Minimizing Side Effects

  • 设计新方法让技能发现聚焦特定状态变量
  • 状态空间覆盖提升三倍,学习效率显著提高
  • 适合需要精确控制的强化学习下游任务

技能对解决复杂问题至关重要。现有技能发现算法常忽略许多强化学习任务中固有的状态变量,导致发现的技能无法精确控制特定变量,严重影响探索效率,增加下游任务中的负面副作用。本文提出一种通用方法,使技能发现算法能够学习到聚焦技能——即专门控制特定状态变量的技能。该方法使状态空间覆盖提升三倍,解锁新的学习能力,并在目标不明确的下游任务中自动规避负面副作用。

原文摘要 · Abstract (English)

Skills are essential for unlocking higher levels of problem solving. A common approach to discovering these skills is to learn ones that reliably reach different states, thus empowering the agent to control its environment. However, existing skill discovery algorithms often overlook the natural state variables present in many reinforcement learning problems, meaning that the discovered skills lack control of specific state variables. This can significantly hamper exploration efficiency, make skills more challenging to learn with, and lead to negative side effects in downstream tasks when the goal is under-specified. We introduce a general method that enables these skill discovery algorithms to learn focused skills -- skills that target and control specific state variables. Our approach improves state space coverage by a factor of three, unlocks new learning capabilities, and automatically avoids negative side effects in downstream tasks.

技能发现强化学习状态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。