无需精确建模,无人机可自适应提升飞行敏捷性。
Learning Agile Quadrotor Flight in the Real World
- 通过动态时间缩放与在线残差学习,实时更新飞行模型。
- 100秒内将最高速度从1.9米/秒提升至7.3米/秒。
- 适合需要高机动性且环境多变的现实飞行场景。
基于学习的控制器在敏捷无人机飞行中表现优异,但通常依赖大量仿真训练,需精确系统辨识以实现仿真到现实的迁移。然而,即使建模精准,固定策略仍易受分布外扰动影响,如气动干扰或硬件退化,迫使控制器采取保守安全策略,限制了实际环境中的机动能力。在线适应虽为解决方案,但物理极限探索受限于数据稀缺与安全风险。为此,我们提出自适应框架,无需精确建模或离线仿真实现。引入自适应时间缩放(ATS)主动探索平台物理极限,并采用在线残差学习增强简单基准模型。基于学习的混合模型,进一步提出现实锚定短时程反向传播(RASH-BPTT),实现高效鲁棒的机上策略更新。大量实验表明,该系统能可靠执行接近执行器饱和极限的敏捷动作。系统在约100秒飞行时间内,将基础策略速度从1.9米/秒提升至7.3米/秒。结果表明,现实自适应不仅是补偿建模误差的手段,更是实现激进飞行模式下持续性能优化的实际机制。
原文摘要 · Abstract (English)
Learning-based controllers have achieved impressive performance in agile quadrotor flight but typically rely on massive training in simulation, necessitating accurate system identification for effective Sim2Real transfer. However, even with precise modeling, fixed policies remain susceptible to out-of-distribution scenarios, ranging from external aerodynamic disturbances to internal hardware degradation. To ensure safety under these evolving uncertainties, such controllers are forced to operate with conservative safety margins, inherently constraining their agility outside of controlled settings. While online adaptation offers a potential remedy, safely exploring physical limits remains a critical bottleneck due to data scarcity and safety risks. To bridge this gap, we propose a self-adaptive framework that eliminates the need for precise system identification or offline Sim2Real transfer. We introduce Adaptive Temporal Scaling (ATS) to actively explore platform physical limits, and employ online residual learning to augment a simple nominal model. {Based on the learned hybrid model, we further propose Real-world Anchored Short-horizon Backpropagation Through Time (RASH-BPTT) to achieve efficient and robust in-flight policy updates. Extensive experiments demonstrate that our quadrotor reliably executes agile maneuvers near actuator saturation limits. The system evolves a conservative base policy with a peak speed of 1.9 m/s to 7.3 m/s within approximately 100 seconds of flight time. These findings underscore that real-world adaptation serves not merely to compensate for modeling errors, but as a practical mechanism for sustained performance improvement in aggressive flight regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。