arXiv:2503.10419eess.SYcs.RO2025-03被引 2

用深度强化学习实现6自由度运动模拟的实时高精度轨迹规划。

A nonlinear real time capable motion cueing algorithm based on deep reinforcement learning

  • 基于深度强化学习构建非线性运动映射算法,融合完整运动平台动力学模型。
  • 在6自由度系统中实现与传统方法相当的性能,且满足实时性要求。
  • 适合需要高保真动态模拟的飞行/驾驶仿真场景,尤其适用于复杂非线性平台。

在运动仿真中,运动映射算法用于规划运动模拟器平台的轨迹,但工作空间限制导致无法直接复现参考轨迹。运动洗出策略(如将平台返回中心)在此类场景中至关重要。对于具有高度非线性工作空间的串联式多自由度运动平台(MSP),最大化其运动学与动力学能力的利用效率极为关键。传统方法如经典洗出滤波和线性模型预测控制未能考虑平台特有的非线性特性;而尽管非线性模型预测控制更为全面,但其计算开销过大,难以在无需简化的情况下实现实时、有人在环的应用。为克服这些局限,本文首次提出一种基于深度强化学习(DRL)的运动映射算法,并应用于6自由度设置,充分考虑了平台的运动学非线性。此前作者已成功将DRL应用于简化的2自由度系统,未考虑运动学与动力学约束。本研究将其扩展至全部6自由度,通过在算法中集成完整的运动平台运动学模型,实现对真实运动模拟器的适用性。DRL-MCA的训练采用近端策略优化(PPO)的演员-评论家架构,并结合自动超参数优化。在详细阐述训练框架与算法结构后,我们进行了全面验证,结果表明:所提DRL-MCA在性能上可与现有成熟算法媲美,生成的轨迹均满足所有系统约束,且达到实时性要求,计算延迟极低。

原文摘要 · Abstract (English)

In motion simulation, motion cueing algorithms are used for the trajectory planning of the motion simulator platform, where workspace limitations prevent direct reproduction of reference trajectories. Strategies such as motion washout, which return the platform to its center, are crucial in these settings. For serial robotic MSPs with highly nonlinear workspaces, it is essential to maximize the efficient utilization of the MSPs kinematic and dynamic capabilities. Traditional approaches, including classical washout filtering and linear model predictive control, fail to consider platform-specific, nonlinear properties, while nonlinear model predictive control, though comprehensive, imposes high computational demands that hinder real-time, pilot-in-the-loop application without further simplification. To overcome these limitations, we introduce a novel approach using deep reinforcement learning for motion cueing, demonstrated here for the first time in a 6-degree-of-freedom setting with full consideration of the MSPs kinematic nonlinearities. Previous work by the authors successfully demonstrated the application of DRL to a simplified 2-DOF setup, which did not consider kinematic or dynamic constraints. This approach has been extended to all 6 DOF by incorporating a complete kinematic model of the MSP into the algorithm, a crucial step for enabling its application on a real motion simulator. The training of the DRL-MCA is based on Proximal Policy Optimization in an actor-critic implementation combined with an automated hyperparameter optimization. After detailing the necessary training framework and the algorithm itself, we provide a comprehensive validation, demonstrating that the DRL MCA achieves competitive performance against established algorithms. Moreover, it generates feasible trajectories by respecting all system constraints and meets all real-time requirements with low...

运动模拟强化学习实时控制6自由度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。