arXiv:2606.27603cs.RO2026-06

用强化学习让无人机精准抛投货物,误差减半、速度提升30%。

Learning to Throw: Agile and Accurate Cable-Suspended Payload Delivery with a Quadrotor

论文配图:Learning to Throw: Agile and Accurate Cable-Suspended Payload Delivery with a Quadrotor
图 1 · 摘自论文原文
  • 融合物理引擎与无人机模型,实现高精度悬吊载荷仿真
  • 训练的智能体使抛投误差降低50%,时长缩短30%
  • 支持视觉输入,适合真实飞行场景快速部署

四旋翼无人机在搜救、医疗物资投送等紧急任务中具备快速运输悬吊载荷的机动性。尽管悬吊运输已有广泛研究,但高动态目标释放仍较薄弱。现有方法依赖基于模型的轨迹优化与跟踪,常因保守约束、跟踪误差及柔性绳索动力学难以解析建模而表现不佳。为此,我们提出一种混合仿真框架,将高保真四旋翼模型与复杂绳索-载荷交互的物理求解器耦合,在每一步交换受力,实现悬吊系统精确模拟。基于此环境,训练了一个深度强化学习策略,可执行敏捷且精准的载荷抛投。该策略零样本部署于硬件,性能超越基于模型的基线:落地误差最多减少50%,抛投时长最多缩短30%。消融实验表明,耦合仿真为性能提升关键。此外,同一流程还可训练仅依赖视觉观测的策略,其精度与状态感知策略相当。为推动动态空中操作研究,我们在论文接受后开源仿真器。

原文摘要 · Abstract (English)

Quadrotors offer the agility needed to rapidly transport suspended payloads during time-critical applications, including search-and-rescue and medical delivery. While suspended-payload transport and traversal for these missions are well studied, the highly dynamic targeted release of the payload remains comparatively underexplored. State-of-the-art approaches typically rely on model-based trajectory optimization and tracking; however, these methods often yield sub-optimal performance due to conservative feasibility constraints, tracking errors, and the inherent difficulty of analytically modeling flexible rope dynamics. To overcome these limitations, we propose a hybrid simulation framework that couples a high-fidelity analytical quadrotor model with a physics solver for complex rope and payload interactions. By exchanging forces between the two domains at every step, we obtain a physically accurate simulation of the suspended-payload system. Leveraging this environment, we train a deep reinforcement learning (RL) policy that executes agile, accurate payload throws to designated targets. Deployed zero-shot on hardware, our RL policy pushes the boundary of the agility-accuracy trade-off, outperforming the model-based baseline by reducing the landing error by up to 50% and the throw duration by up to 30%. Ablation studies confirm that the coupled simulation is the key enabler of these gains. We further show that the same pipeline trains a policy driven by visual observations rather than an explicit state estimate, achieving accuracy comparable to that of the state-based policy. To accelerate future research in dynamic aerial manipulation, we open-source the simulator to the community upon acceptance.

无人机强化学习动态投送物理仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。