arXiv:2607.25985cs.ROcs.LG2026-07

基于物理模型的强化学习,直接控制四旋翼低层力矩与推力。

Physics-Aware End-to-End Deep Reinforcement Learning for Quadcopter Control with Actuator Dynamics

  • 在高保真Simulink环境中,直接输出推力与力矩控制四旋翼
  • SAC和TD3算法在稳定性和探索效率上表现最优,电机延迟影响显著
  • 适合研究无人机低层控制或强化学习与物理建模结合的学者

无人飞行器(尤其是四旋翼)因欠驱动特性面临独特控制挑战:仅四个控制输入需调控六个自由度。本文研究一种融合物理先验的端到端深度强化学习方法,直接作用于低层体控输入——总推力与机体力矩(T, τ_x, τ_y, τ_z),并通过高保真Simulink环境闭环训练。仿真器集成12状态刚体模型(MATLAB Level-2 S-Function),包含:(i) 基于摩尔-彭罗斯伪逆的力矩分配(由推力与阻力项导出系数矩阵),(ii) 每个电机的一阶动态(时间常数 $T_m = 0.076$ s),含转子陀螺耦合效应。奖励函数通过指数位置势阱、姿态惩罚及速度二次代价平衡目标到达与稳定性。评估四种算法(DDPG、TD3、PPO、SAC)在两阶段任务中表现:(S1) 仅推力悬停;(S2) 含俯仰力矩及偏移目标悬停。结果表明,SAC与TD3在稳定性与探索效率上更优,而PPO样本效率较低。研究强调了建模电机滞后与气动力矩对低层控制稳定性的关键作用,并提供可复现的四旋翼强化学习基准。

原文摘要 · Abstract (English)

Unmanned aerial vehicles (UAVs), particularly quadcopters, present unique challenges for autonomous control due to their underactuated dynamics: only four available control inputs must govern six degrees of freedom. This paper investigates a physics-aware, end-to-end deep reinforcement learning (DRL) approach that acts directly on low-level body inputs, total thrust and body torques $(T, τ_x, τ_y, τ_z)$, and closes the loop through a high-fidelity Simulink environment. Our simulator integrates a 12-state rigid-body model (MATLAB Level-2 S-Function) with (i) an Action2RPM allocation based on the Moore-Penrose pseudo-inverse of a coefficient matrix derived from thrust and drag terms, and (ii) first-order actuator dynamics for each motor (time constant $T_m = 0.076$ s), including rotor gyroscopic coupling. A shaped reward balances goal-reaching and stability using an exponential position well, attitude penalties, and quadratic velocity costs. Four DRL algorithms, DDPG, TD3, PPO, and SAC, are evaluated in two stages: (S1) thrust-only hover and (S2) hover with pitch torque and a translated goal. Results show that SAC and TD3 achieve superior stability and exploration efficiency, while PPO is less sample-efficient. The study highlights the significance of modeling actuator lags and aerodynamic moments for stable low-level control and provides a reproducible benchmark for quadcopter DRL.

强化学习无人机控制物理建模端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。