arXiv:2504.01915cs.NEcs.RO2025-04被引 6

无监督质量多样性算法让机器人控制优化摆脱陷阱,无需人工设计特征

Overcoming Deceptiveness in Fitness Optimization with Unsupervised Quality-Diversity

  • 用自监督学习自动提取感官数据特征,替代人工设计多样性度量
  • 在多个任务上超越传统优化方法,最高提升34%性能
  • 适合缺乏先验知识的复杂控制场景,如未知环境下的机器人训练

策略优化根据目标函数寻找控制问题的最优解,是工程与研究的基础领域,广泛应用于机器人。传统方法如强化学习和进化算法在具有欺骗性收益曲面的问题中表现不佳,因局部最优会误导搜索方向。质量多样性(QD)算法通过维持多样化的中间解作为跳出局部极值的跳板,提供新思路。但现有QD算法需依赖领域专家设计手工特征,限制了其在难以定义解多样性的情形下的应用。本文提出AURORA-XCon框架,结合对比学习与周期性灭绝事件,使无监督QD算法能从感官数据中自动学习特征,在无需领域知识的情况下高效解决欺骗性优化问题。实验显示,该方法优于所有传统优化基线,并在部分任务上超过使用领域特化特征的最佳QD基线,最高提升达34%。本工作将无监督QD的应用从发现新颖解转向传统优化,拓展了其在特征空间难定义领域的潜力。

原文摘要 · Abstract (English)

Policy optimization seeks the best solution to a control problem according to an objective or fitness function, serving as a fundamental field of engineering and research with applications in robotics. Traditional optimization methods like reinforcement learning and evolutionary algorithms struggle with deceptive fitness landscapes, where following immediate improvements leads to suboptimal solutions. Quality-diversity (QD) algorithms offer a promising approach by maintaining diverse intermediate solutions as stepping stones for escaping local optima. However, QD algorithms require domain expertise to define hand-crafted features, limiting their applicability where characterizing solution diversity remains unclear. In this paper, we show that unsupervised QD algorithms - specifically the AURORA framework, which learns features from sensory data - efficiently solve deceptive optimization problems without domain expertise. By enhancing AURORA with contrastive learning and periodic extinction events, we propose AURORA-XCon, which outperforms all traditional optimization baselines and matches, in some cases even improving by up to 34%, the best QD baseline with domain-specific hand-crafted features. This work establishes a novel application of unsupervised QD algorithms, shifting their focus from discovering novel solutions toward traditional optimization and expanding their potential to domains where defining feature spaces poses challenges.

强化学习质量多样性无监督学习机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。