针对无人机视觉导航中误检问题,提出可靠性感知的粗粒度目标优化方法。
RACO: Reliability-Aware Coarse-Goal Optimization for Inspection-Oriented UAV Vision-Language Navigation

- 将粗略目标视为动态假设,用候选锚点实时验证并修正定位
- 在未见数据上提升成功率9.53和7.98个百分点,降低误确认风险
- 适合需要精准停靠与避错的巡检类无人机应用
无人机视觉语言导航(UAV-VLN)通常以到达目标点为评估标准,但在巡检场景中,需确保智能体在有效检查区域内停止,并避免误认视觉或语义相似的干扰项。现有粗到细导航策略的关键缺陷在于:粗粒度目标被默认可靠,但可能漂移至看似合理却错误的物体区域,限制局部阶段的修复能力。为此,我们引入LG-UVI——基于CityNav/CityRefer的对象中心巡检评估设置,包含目标物体、强干扰项、类型感知检查区及检查区到达与对象级确认的诊断机制。为应对该场景,提出RACO:一种可靠性感知的自适应粗到细导航框架。RACO不将预测粗粒度目标当作固定路径点,而是视作运行时假设,在阶段一前及阶段一至阶段二边界,利用对象级候选锚点进行校验与修正;同时采用尺度自适应终端精修,结合运行时可观察的几何与锚点证据处理终端近似情况。在统一在线评估协议下,RACO相较复现的HETT基线,在验证集未见和测试集未见数据上分别提升成功率达9.53和7.98个百分点,同时提高检查区到达率并降低误验证风险,表明粗粒度目标可靠性优化是现有策略的有效补充。
原文摘要 · Abstract (English)
UAV vision-language navigation (UAV-VLN) is commonly evaluated as goal reaching, but inspection-oriented deployment requires the agent to stop within a valid inspection region and avoid falsely confirming visually or semantically similar distractors. This requirement exposes a key weakness in existing coarse-to-fine UAV-VLN policies: the coarse goal predicted before local refinement is often treated as reliable, although it may drift toward plausible but incorrect object regions and limit the ability of the local stage to recover. To systematically evaluate this problem, we introduce LG-UVI, an object-centric inspection evaluation setting derived from CityNav/CityRefer. LG-UVI extends standard UAV-VLN episodes with target objects, hard distractors, type-aware inspection regions, and diagnostics for inspection-region arrival and object-level confirmation. To address this inspection-oriented setting, we further propose RACO, a reliability-aware adaptive coarse-to-fine navigation framework. Instead of treating the predicted coarse goal as a fixed waypoint, RACO views it as a runtime hypothesis and uses object-level candidate anchors to check and correct coarse localization before Stage 1 and at the Stage 1-to-Stage 2 boundary. RACO also applies scale-adaptive terminal refinement to handle terminal near-miss cases using runtime-observable geometric and anchor-based evidence. Under a unified online evaluation protocol, RACO improves SR over the reproduced HETT baseline by 9.53 and 7.98 percentage points on validation-unseen and test-unseen, respectively. It also improves inspection-region arrival and reduces false verification risk, showing that coarse-goal reliability optimization is an effective complement to existing coarse-to-fine UAV-VLN policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。