arXiv:2510.05996cs.AIcs.IT2025-10被引 1

用信息论方法预训练智能体,提升下游任务的数据效率。

Information-Theoretic Policy Pre-Training with Empowerment

  • 引入折扣激励度,平衡短期与长期控制能力。
  • 最大化折扣激励度的策略在下游任务中数据效率更高。
  • 适合需要快速适应的新任务,尤其适用于强化学习初学者。

激励度(Empowerment)是一种衡量智能体对环境潜在影响的信息论指标,已成为强化学习中强大的内在动机与探索框架。尽管已在无监督强化学习和技能学习中应用,但将其作为预训练信号的研究仍较有限。本文表明,激励度可作为数据高效下游任务适配的预训练信号。为此,我们扩展了传统激励度概念,提出折扣激励度,以在短中期与长期视野间平衡智能体对环境的控制能力。基于此,我们提出一种新预训练范式:初始化策略以最大化折扣激励度,使智能体获得对环境动态的稳健理解。我们分析了该方法在多种现有强化学习算法中的表现,并通过实验证明其具备通用初始化潜力:具有长时域的激励度最大化策略在下游任务中表现出更高的数据效率与适应性。研究为未来将该框架扩展至高维复杂任务铺平道路,进一步推动强化学习发展。

原文摘要 · Abstract (English)

Empowerment, an information-theoretic measure of an agent's potential influence on its environment, has emerged as a powerful intrinsic motivation and exploration framework for reinforcement learning (RL). Besides for unsupervised RL and skill learning algorithms, the specific use of empowerment as a pre-training signal has received limited attention in the literature. We show that empowerment can be used as a pre-training signal for data-efficient downstream task adaptation. For this we extend the traditional notion of empowerment by introducing discounted empowerment, which balances the agent's control over the environment across short- and long-term horizons. Leveraging this formulation, we propose a novel pre-training paradigm that initializes policies to maximize discounted empowerment, enabling agents to acquire a robust understanding of environmental dynamics. We analyze empowerment-based pre-training for various existing RL algorithms and empirically demonstrate its potential as a general-purpose initialization strategy: empowerment-maximizing policies with long horizons are data-efficient and effective, leading to improved adaptability in downstream tasks. Our findings pave the way for future research to scale this framework to high-dimensional and complex tasks, further advancing the field of RL.

强化学习激励度预训练数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。