arXiv:2509.02930cs.LGcs.AI2025-09被引 2

用生态多样性指标提升强化学习技能多样性,支持多场景预训练。

VendiRL: A Framework for Self-Supervised Reinforcement Learning of Diversely Diverse Skills

  • 引入生态学中的Vendi Score衡量技能多样性,灵活定义多样形式。
  • 在多种相似性函数下实现多样化技能学习,适配不同任务需求。
  • 解决多样性评估不统一问题,适合复杂交互环境的技能预训练。

自监督强化学习中,如何学习多样技能以应对未知未来任务是关键挑战。现有方法在可扩展性和评估方面仍存不足:高维特征空间中相关特征可能随下游任务变化;多样性定义依赖特定假设,导致评估标准不一致,难以横向比较,且多数多样性形式未被探索。为此,本文采用源自生态学的样本多样性度量——Vendi Score,使用户可灵活指定和评估任意形式的多样性。该度量促进技能评估,并推动构建统一框架VendiRL,通过不同相似性函数驱动多样化技能生成,适用于新奇且高度互动的环境,支持针对多种多样性目标的预训练。

原文摘要 · Abstract (English)

In self-supervised reinforcement learning (RL), one of the key challenges is learning a diverse set of skills to prepare agents for unknown future tasks. Despite impressive advances, scalability and evaluation remain prevalent issues. Regarding scalability, the search for meaningful skills can be obscured by high-dimensional feature spaces, where relevant features may vary across downstream task domains. For evaluating skill diversity, defining what constitutes "diversity" typically requires a hard commitment to a specific notion of what it means for skills to be diverse, potentially leading to inconsistencies in how skill diversity is understood, making results across different approaches hard to compare, and leaving many forms of diversity unexplored. To address these issues, we adopt a measure of sample diversity that translates ideas from ecology to machine learning -- the Vendi Score -- allowing the user to specify and evaluate any desired form of diversity. We demonstrate how this metric facilitates skill evaluation and introduce VendiRL, a unified framework for learning diversely diverse sets of skills. Given distinct similarity functions, VendiRL motivates distinct forms of diversity, which could support skill-diversity pretraining in new and richly interactive environments where optimising for various forms of diversity may be desirable.

强化学习技能多样性自监督预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。