arXiv:2603.08019cs.RO2026-03

用向量场增强梯度信号,让无人机竞速更稳更快

Vector Field Augmented Differentiable Policy Learning for Vision-Based Drone Racing

  • 引入向量场提供连续梯度,优化穿越门和避障的平衡
  • 在仿真和真实环境中实现更快收敛与更高飞行鲁棒性
  • 无需系统辨识即可高效实现实验室到真实场景迁移

自主无人机在复杂环境中竞速需要高速敏捷飞行同时保持可靠的障碍物规避。基于可微物理的策略学习方法近期在多种任务中展现出高样本效率和优异性能,包括高速无人机飞行和四足机器人运动。然而,将此类方法应用于无人机竞速仍面临挑战,因为关键目标如穿越门难以表达为平滑、可微的损失函数。为此,我们提出DiffRacing,一种新型的向量场增强型可微策略学习框架。该框架将可微损失与向量场结合,在训练过程中提供连续且稳定的梯度信号,平衡障碍物规避与高速穿越门的需求。此外,引入可微的Delta动作模型以补偿动力学差异,实现无需显式系统辨识的高效仿真实现到真实场景的迁移。大量仿真与真实世界实验表明,DiffRacing在样本效率、收敛速度和飞行鲁棒性方面均表现卓越,证明了向量场能为传统基于梯度的策略学习引入任务相关的几何先验。

原文摘要 · Abstract (English)

Autonomous drone racing in complex environments requires agile, high-speed flight while maintaining reliable obstacle avoidance. Differentiable-physics-based policy learning has recently demonstrated high sample efficiency and remarkable performance across various tasks, including agile drone flight and quadruped locomotion. However, applying such methods to drone racing remains difficult, as key objective like gate traversal are inherently hard to express as smooth, differentiable losses. To address these challenges, we propose DiffRacing, a novel vector field-augmented differentiable policy learning framework. DiffRacing integrates differentiable losses and vector fields into the training process to provide continuous and stable gradient signals, balancing obstacle avoidance and high-speed gate traversal. In addition, a differentiable Delta Action Model compensates for dynamics mismatch, enabling efficient sim-to-real transfer without explicit system identification. Extensive simulation and real-world experiments demonstrate that DiffRacing achieves superior sample efficiency, faster convergence, and robust flight performance, thereby demonstrating that vector fields can augment traditional gradient-based policy learning with a task-specific geometric prior.

无人机竞速可微物理向量场策略学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。