用相似性度量优化控制模型,提升稳定性和效率
Bisimulation metric for Model Predictive Control
- 在目标函数中加入双仿真度量损失,直接优化编码器
- 训练时间减少,对输入噪声更鲁棒,稳定性显著提升
- 适合需要高效可靠决策的复杂控制场景
基于模型的强化学习在复杂环境中的样本效率和决策能力方面展现出潜力。然而,现有方法在训练稳定性、抗噪声能力及计算效率方面仍面临挑战。本文提出双仿真度量模型预测控制(BS-MPC),在目标函数中引入双仿真度量损失,直接优化编码器。这种逐时间步的直接优化使学习到的编码器能从原始状态空间中提取内在信息,同时剔除无关细节,并防止梯度与误差发散。BS-MPC通过减少训练时间,提升了训练稳定性、对输入噪声的鲁棒性及计算效率。我们在DeepMind Control Suite的连续控制和基于图像的任务上进行了评估,结果表明其性能和鲁棒性优于当前最优基线方法。
原文摘要 · Abstract (English)
Model-based reinforcement learning has shown promise for improving sample efficiency and decision-making in complex environments. However, existing methods face challenges in training stability, robustness to noise, and computational efficiency. In this paper, we propose Bisimulation Metric for Model Predictive Control (BS-MPC), a novel approach that incorporates bisimulation metric loss in its objective function to directly optimize the encoder. This time-step-wise direct optimization enables the learned encoder to extract intrinsic information from the original state space while discarding irrelevant details and preventing the gradients and errors from diverging. BS-MPC improves training stability, robustness against input noise, and computational efficiency by reducing training time. We evaluate BS-MPC on both continuous control and image-based tasks from the DeepMind Control Suite, demonstrating superior performance and robustness compared to state-of-the-art baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。