arXiv:2512.10934cs.ROcs.LG2025-12

用课程学习让无人机在未知管道中自主导航,不依赖几何信息也能稳定前行。

Curriculum-Based Reinforcement Learning for Autonomous UAV Navigation in Unknown Curved Tubular Conduit

  • 通过渐进式课程学习,逐步增加管道弯曲难度训练智能体。
  • 在真实3D环境中,强化学习策略性能超越已知路径的纯追踪算法。
  • 适合工业巡检、地下或医疗内窥等狭窄空间自主导航场景。

无人机在受限管状环境中的自主导航仍面临巨大挑战,主要源于管道几何约束、靠近壁面以及感知局限。本文提出一种基于强化学习的方法,使无人机在完全未知的三维管道中导航,仅依赖激光雷达(LiDAR)的局部观测和对管道中心的条件性视觉检测,无需任何先验几何知识。与可获取中心线的确定性纯追踪算法相比,该方法设计了信息不对称以评估强化学习在缺乏几何模型时的补偿能力。智能体采用渐进式课程学习策略,逐步暴露于曲率递增的管道环境,此时管道中心频繁脱离视觉视野。结合直接可见性、方向记忆和激光雷达对称性线索的转向协商机制,在部分可观测条件下确保了导航稳定性。实验表明,采用PPO算法的策略展现出鲁棒且可泛化的行为,即使在信息受限情况下仍持续优于确定性控制器。高保真3D环境验证了所学策略向连续物理动力学系统的迁移能力。该方法为未知管状环境下的自主导航提供了完整框架,并为工业、地下或医疗应用中穿越狭窄、弱感知通道开辟了新路径。

原文摘要 · Abstract (English)

Autonomous drone navigation in confined tubular environments remains a major challenge due to the constraining geometry of the conduits, the proximity of the walls, and the perceptual limitations inherent to such scenarios. We propose a reinforcement learning approach enabling a drone to navigate unknown three-dimensional tubes without any prior knowledge of their geometry, relying solely on local observations from LiDAR and a conditional visual detection of the tube center. In contrast, the Pure Pursuit algorithm, used as a deterministic baseline, benefits from explicit access to the centerline, creating an information asymmetry designed to assess the ability of RL to compensate for the absence of a geometric model. The agent is trained through a progressive Curriculum Learning strategy that gradually exposes it to increasingly curved geometries, where the tube center frequently disappears from the visual field. A turning-negotiation mechanism, based on the combination of direct visibility, directional memory, and LiDAR symmetry cues, proves essential for ensuring stable navigation under such partial observability conditions. Experiments show that the PPO policy acquires robust and generalizable behavior, consistently outperforming the deterministic controller despite its limited access to geometric information. Validation in a high-fidelity 3D environment further confirms the transferability of the learned behavior to a continuous physical dynamics. The proposed approach thus provides a complete framework for autonomous navigation in unknown tubular environments and opens perspectives for industrial, underground, or medical applications where progressing through narrow and weakly perceptive conduits represents a central challenge.

无人机导航强化学习管道巡检课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。