让无人机通过看图找物,智能选视角,安全飞向目标
Flying to Image-Specified Objects: 3D Quadrotor Navigation via Cross-Graph Memory and Viewpoint Planning

- 分层导航:先选好视角再飞行,提升定位精度
- 轻量记忆模块让系统记住物体和观察信息,辅助决策
- 真实飞行测试验证方案有效,适合复杂环境自主导航
实例指定图像目标导航(InstanceImageNav)要求机器人精准抵达查询图像中指定的物体实例。将此任务扩展至四旋翼无人机面临连续三维控制、视场受限及安全约束等挑战,成功导航高度依赖于选择具有信息量的观测视角。本文提出一种分层导航框架,将高层决策与底层运动执行分离。系统不在空间位置上直接导航,而是围绕前沿区域和潜在目标物体生成感知导向的动作节点,使无人机在探索过程中保持对目标实例的可观测性。设计轻量级语义记忆,维护物体级别与观测级别的上下文信息,实现语义线索向候选动作节点的传播,支持更优决策。基于学习的策略选择最具前景的动作节点,轨迹规划器生成动态可行的三维飞行路径以确保安全执行。仿真实验显示性能持续优于强基线方法,真实四旋翼飞行测试验证了该框架的实用性与鲁棒性。
原文摘要 · Abstract (English)
Instance-Specific Image-Goal Navigation (InstanceImageNav) requires a robot to navigate toward the exact object instance depicted in a query image. Extending this task to quadrotors is challenging due to continuous 3D control, limited field of view (FOV), and safety constraints, which make successful navigation highly dependent on selecting informative viewpoints. We propose a hierarchical navigation framework for quadrotor InstanceImageNav that separates high-level decision making from low-level motion execution. Instead of navigating directly to spatial locations, the system generates viewpoint-aware action nodes around frontier regions and potential target objects, enabling the robot to explore while maintaining informative viewpoints for detecting the target instance. A lightweight semantic memory maintains object-level and observation-level context, allowing semantic cues to propagate to candidate action nodes for decision making. A learning-based policy selects the most promising action node, and a trajectory planner generates dynamically feasible 3D flight paths for safe execution. Experiments in simulation demonstrate consistent improvements over strong baselines, and real-world quadrotor flights validate the practicality and robustness of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。