让无人机通过视觉语言模型更准地找目标,提升导航可靠性。
Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

- 分两层构建语义与目标证据的持久空间信念图
- 在无人机基准上达成21.61%成功率、35.57%目标寻得率
- 适合做空中目标导航的算法研究者和工程应用
航拍目标导航(ObjectNav)要求无人飞行器(UAV)在未知户外环境中,仅依靠机载视觉观测定位指定目标。视觉语言模型(VLM)可理解开放式目标描述与视觉信息,但其帧级输出常具噪声、稀疏且空间瞬时。本文提出AeroBelief框架,将临时的VLM观测转化为持久的空间引导:通过直觉层积累场景级语义线索用于探索,证据层保留合格的目标特定观测用于接近与确认;采用证据门控融合生成空间信念热点。进一步引入对象条件下的保守证据判定机制,提升观测可靠性。同时,以本体为中心的区域引导将四叉树覆盖转为航向对齐的方向性提议,并通过时间承诺稳定策略,其区域评分独立于语义信念值,维持探索压力并减少低收益重复搜索。在UAV-ON基准上的实验表明,AeroBelief在所有对比方法中达到最优综合表现,成功率(SR)、目标寻得率(OSR)与路径相似度(SPL)分别达21.61%、35.57%和10.62。结果验证了持久语义-空间信念、保守证据判定与时间稳定区域引导的有效性。
原文摘要 · Abstract (English)
Aerial Object Goal Navigation (ObjectNav) requires an unmanned aerial vehicle (UAV) to locate a described target in an unknown outdoor environment using onboard visual observations. Vision-language models (VLMs) can interpret open-ended target descriptions and visual observations, but their frame-level outputs are often noisy, sparse, and spatially transient. We propose AeroBelief, a dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance. It separates broad contextual plausibility from target-specific evidence: an intuition layer accumulates scene-level semantic cues for exploration, while an evidence layer preserves qualified target-specific observations for approach and confirmation. Evidence-gated fusion combines the two layers into spatial belief hotspots. We further introduce object-conditioned visual reasoning with conservative evidence qualification to improve observation reliability before spatial accumulation. In parallel, egocentric regional guidance converts quadtree coverage into UAV-centered, yaw-aligned directional proposals and stabilizes them through temporal commitment. Its regional scoring is independent of semantic belief values, maintaining exploration pressure and reducing repeated low-gain search. Experiments on the UAV-ON benchmark show that AeroBelief achieves the best reported overall SR, OSR, and SPL among the compared methods, reaching 21.61%, 35.57%, and 10.62, respectively. These results support the effectiveness of persistent semantic-spatial belief, conservative evidence qualification, and temporally stable regional guidance for aerial ObjectNav.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。