arXiv:2508.18817cs.ROcs.LG2025-08被引 1

用偏好学习让无人机学会高难度特技飞行,效果接近人工设计奖励。

Learning Acrobatic Flight from Preferences

  • 基于偏好构建奖励模型的集成框架,显式建模每一步的奖励不确定性。
  • 在四轴飞行器特技控制中达到人工奖励88.4%的性能,远超传统方法。
  • 仅靠人类偏好反馈训练,可直接部署到真实无人机上执行复杂动作。

基于偏好的强化学习(PbRL)使智能体无需人工设计奖励函数即可学习控制策略,特别适用于目标难以形式化或主观性强的任务。特技飞行因动态复杂、动作迅速且执行精度要求高而极具挑战性。我们发现,人工设计的奖励与人类判断一致率仅为60.7%,凸显了偏好驱动方法的必要性。本文提出奖励集成置信度(REC)框架,通过分布式奖励模型集合显式建模每一步的奖励不确定性,并将不确定性传播至偏好损失,利用分歧促进探索。在四轴飞行器特技控制任务中,REC性能达人工奖励的88.4%,显著优于标准偏好PPO的55.2%。我们在仿真中训练策略并实现零样本迁移至真实世界,成功展示纯偏好反馈学习的复杂特技动作。此外,我们在连续控制基准测试中验证了REC的跨领域适用性。

原文摘要 · Abstract (English)

Preference-based reinforcement learning (PbRL) enables agents to learn control policies without requiring manually designed reward functions, making it well-suited for tasks where objectives are difficult to formalize or inherently subjective. Acrobatic flight poses a particularly challenging problem due to its complex dynamics, rapid movements, and the importance of precise execution. However, manually designed reward functions for such tasks often fail to capture the qualities that matter: we find that hand-crafted rewards agree with human judgment only 60.7% of the time, underscoring the need for preference-driven approaches. In this work, we propose Reward Ensemble under Confidence (REC), a probabilistic reward learning framework for PbRL that explicitly models per-timestep reward uncertainty through an ensemble of distributional reward models. By propagating uncertainty into the preference loss and leveraging disagreement for exploration, REC achieves 88.4% of shaped reward performance on acrobatic quadrotor control, compared to 55.2% with standard Preference PPO. We train policies in simulation and successfully transfer them zero-shot to the real world, demonstrating complex acrobatic maneuvers learned purely from preference feedback. We further validate REC on a continuous control benchmark, confirming its applicability beyond the domain of aerial robotics.

偏好学习强化学习无人机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。