arXiv:2604.12916cs.RO2026-04

E2E-Fly让四旋翼无人机从仿真到真实世界一键部署,无需重新训练。

E2E-Fly: An Integrated Training-to-Deployment System for End-to-End Quadrotor Autonomy

论文配图:E2E-Fly: An Integrated Training-to-Deployment System for End-to-End Quadrotor Autonomy
图 1 · 摘自论文原文
  • 融合可微分物理的仿真平台,支持端到端强化学习训练。
  • 六类控制任务在真实无人机上成功部署,实现零样本迁移。
  • 提供从训练到部署的完整流程,适合机器人研发与工业应用。

由于视觉渲染效率低、物理建模不准确、传感器差异未建模,以及缺乏集成可微分物理学习的统一训练-部署平台,基于学习的四旋翼飞行器策略从仿真迁移到现实仍具挑战。尽管近期研究已实现多种端到端控制任务,但系统性、零样本迁移的全流程框架仍稀缺,制约了复现与实际部署。为此,我们提出E2E-Fly,一个集成的四旋翼平台与全栈训练、验证、部署工作流。训练框架包含高性能仿真器,支持可微分物理学习与强化学习,并设计结构化奖励函数以适配常见四旋翼任务。我们进一步引入两阶段验证策略:仿真到仿真的迁移测试与硬件在环测试。通过专用底层控制接口与全面的仿真-现实对齐方法(包括系统辨识、域随机化、延迟补偿、噪声建模),将策略部署至两个真实四旋翼平台。据我们所知,这是首个系统性整合可微分物理学习与训练、验证、真实部署的四旋翼工作。最终,我们在真实世界中成功训练并部署了六种端到端控制任务。

原文摘要 · Abstract (English)

Training and transferring learning-based policies for quadrotors from simulation to reality remains challenging due to inefficient visual rendering, physical modeling inaccuracies, unmodeled sensor discrepancies, and the absence of a unified platform integrating differentiable physics learning into end-to-end training. While recent work has demonstrated various end-to-end quadrotor control tasks, few systems provide a systematic, zero-shot transfer pipeline, hindering reproducibility and real-world deployment. To bridge this gap, we introduce E2E-Fly, an integrated framework featuring an agile quadrotor platform coupled with a full-stack training, validation, and deployment workflow. The training framework incorporates a high-performance simulator with support for differentiable physics learning and reinforcement learning, alongside structured reward design tailored to common quadrotor tasks. We further introduce a two-stage validation strategy using sim-to-sim transfer and hardware-in-the-loop testing, and deploy policies onto two physical quadrotor platforms via a dedicated low-level control interface and a comprehensive sim-to-real alignment methodology, encompassing system identification, domain randomization, latency compensation, and noise modeling. To the best of our knowledge, this is the first work to systematically unify differentiable physical learning with training, validation, and real-world deployment for quadrotors. Finally, we demonstrate the effectiveness of our framework for training six end-to-end control tasks and deploy them in the real world.

四旋翼端到端仿真-现实强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。