研究可变结构下信任域算法的优化机制,提升机器人控制策略的泛化能力。
Beyond Fixed Morphologies: Learning Graph Policies with Trust Region Compensation in Variable Action Spaces
- 基于图结构设计可适应不同形态的策略网络
- 发现动作空间维度变化会改变优化稳定性边界
- 适合关注机器人控制泛化与强化学习鲁棒性的研究者
信任区域类优化方法在连续控制任务中表现出稳定性和良好性能,成为强化学习的基础算法。随着对可扩展、可复用控制策略的需求增长,策略对不同运动学结构的泛化能力(即形态泛化)日益重要。图结构策略架构自然且有效地编码了结构差异。然而,尽管此类架构能适应可变形态,信任区域方法在动作空间维度变化时的行为仍不明确。本文针对信任区域策略优化方法,重点分析了TRPO及其广泛应用的一阶近似PPO,揭示动作空间维度变化如何影响优化景观,尤其在KL散度或策略裁剪惩罚约束下的表现。结合理论分析,通过在Gymnasium Swimmer环境中进行实验评估,系统性地改变运动学结构而不改变任务本身,构建了研究形态泛化的理想基准场景。
原文摘要 · Abstract (English)
Trust region-based optimization methods have become foundational reinforcement learning algorithms that offer stability and strong empirical performance in continuous control tasks. Growing interest in scalable and reusable control policies translate also in a demand for morphological generalization, the ability of control policies to cope with different kinematic structures. Graph-based policy architectures provide a natural and effective mechanism to encode such structural differences. However, while these architectures accommodate variable morphologies, the behavior of trust region methods under varying action space dimensionality remains poorly understood. To this end, we conduct a theoretical analysis of trust region-based policy optimization methods, focusing on both Trust Region Policy Optimization (TRPO) and its widely used first-order approximation, Proximal Policy Optimization (PPO). The goal is to demonstrate how varying action space dimensionality influence the optimization landscape, particularly under the constraints imposed by KL-divergence or policy clipping penalties. Complementing the theoretical insights, an empirical evaluation under morphological variation is carried out using the Gymnasium Swimmer environment. This benchmark offers a systematically controlled setting for varying the kinematic structure without altering the underlying task, making it particularly well-suited to study morphological generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。