arXiv:2505.18201cs.ROcs.LG2025-05被引 1

用物理模型与强化学习协同控制扑翼无人机,提升稳定性与采样效率。

Reinforcement Twinning for Hybrid Control of Flapping-Wing Drones

  • 融合物理模型与强化学习,通过数字孪生实时更新模型并共享经验。
  • 三种初始化下均优于纯数据驱动或纯模型方法,样本效率提升显著。
  • 适合需要高鲁棒性与安全性的飞行控制场景,如小型无人机自主飞行。

扑翼无人机的控制需应对时变、非线性、欠驱动的动力学特性,且传感器数据不完整、噪声大。近年来人工智能,特别是强化学习(RL),为解决此类复杂控制问题提供了新思路,通过与环境交互进行数据驱动策略优化。但纯数据驱动方法样本效率低,需大量甚至危险的探索,尤其缺乏物理模型引导时。因此,本研究提出一种混合型无模型/有模型飞行控制框架,采用强化孪生算法。模型基(MB)部分利用伴随公式和自适应数字孪生,从实时轨迹持续识别;无模型(MF)部分使用强化学习。两者通过迁移学习、模仿学习及真实环境与数字孪生间的共享经验协作,由策略裁判根据数字孪生表现和真实-虚拟一致性比率决定哪个代理在现实中执行。该框架在纵向控制中评估,将扑翼无人机建模为由准稳态气动力建驱的非线性时变系统。在三种自适应模型初始化条件下测试:(1)基于已有数据的离线识别,(2)随机初始化并完全在线识别,(3)离线预训练带偏差参数后在线适应。所有情况下,混合框架在性能、鲁棒性和样本效率上均优于纯无模型与纯模型方法。

原文摘要 · Abstract (English)

Controlling flapping-wing drones requires controllers that handle time-varying, nonlinear, underactuated dynamics from incomplete, noisy sensor data. Recent advances in artificial intelligence (AI), particularly reinforcement learning (RL), have opened new perspectives for addressing such complex control problems through data-driven policy optimization from interaction with the environment. Yet purely data-driven methods are sample-inefficient, demanding extensive, sometimes unsafe exploration, especially without guiding physical models. This motivates hybrid AI-physics frameworks. This article proposes a hybrid model-free/model-based flight-control approach using the reinforcement twinning algorithm. The model-based (MB) component uses an adjoint formulation and an adaptive digital twin continuously identified from live trajectories; the model-free (MF) component uses RL. The two agents share knowledge via transfer learning, imitation learning, and shared experience between the real environment and the digital twin, coordinated by a policy referee that selects which agent acts in reality based on digital-twin performance and a real-to-virtual consistency ratio. The framework is evaluated for the longitudinal control of a flapping-wing drone, modelled as a nonlinear time-varying system driven by quasi-steady aerodynamic forces. The hybrid strategy is tested under three adaptive-model initializations: (1) offline identification from existing data, (2) random initialization with fully online identification, and (3) offline pre-training with biased parameters followed by online adaptation. In all cases, the hybrid framework improves performance, robustness, and sample efficiency over purely model-free and purely model-based approaches.

强化学习无人机控制数字孪生混合智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。