仅用单目摄像头实现复杂环境飞行,靠仿真+域适应突破真实世界应用瓶颈。
Flying in Clutter on Monocular RGB by Learning in 3D Radiance Fields with Domain Adaptation
- 在3D高斯泼溅仿真中训练,用对抗域适应缩小仿真与真实差异
- 零样本迁移至物理世界,在不同光照下实现安全敏捷飞行
- 适合做低成本视觉导航的无人机系统研发者参考
现代自主导航系统主要依赖激光雷达和深度相机。然而一个根本性问题仍未解决:仅使用单目RGB图像的飞行机器人能否在复杂环境中导航?由于真实数据采集成本高昂,通过仿真学习政策是一条有前景的路径。但直接将此类策略部署到物理世界时,会受到显著的仿真到现实感知差距阻碍。因此,我们提出一种框架,结合3D高斯泼溅(3DGS)环境的逼真度与对抗域适应技术。通过在高保真仿真中训练,并显式最小化特征差异,我们的方法确保策略依赖于域不变特征。实验结果表明,该策略实现了对物理世界的鲁棒零样本迁移,使机器人能够在光照变化的非结构化环境中实现安全且敏捷的飞行。
原文摘要 · Abstract (English)
Modern autonomous navigation systems predominantly rely on lidar and depth cameras. However, a fundamental question remains: Can flying robots navigate in clutter using solely monocular RGB images? Given the prohibitive costs of real-world data collection, learning policies in simulation offers a promising path. Yet, deploying such policies directly in the physical world is hindered by the significant sim-to-real perception gap. Thus, we propose a framework that couples the photorealism of 3D Gaussian Splatting (3DGS) environments with Adversarial Domain Adaptation. By training in high-fidelity simulation while explicitly minimizing feature discrepancy, our method ensures the policy relies on domain-invariant cues. Experimental results demonstrate that our policy achieves robust zero-shot transfer to the physical world, enabling safe and agile flight in unstructured environments with varying illumination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。