arXiv:2607.25106cs.CV2026-07中稿 · the 2026 IEEE/RSJ …

用网络图片增强文本查询,提升细粒度物体导航准确率

IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation

论文配图:IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation
图 1 · 摘自论文原文
  • 用网页图像丰富文本查询,生成更精准的定位信号
  • 在新基准上导航成功率提升,长尾类别表现显著改善
  • 无需训练模型,适合希望提升导航鲁棒性的研究者

具身智能越来越多依赖预训练视觉-语言模型构建的可查询语义地图,实现零样本物体目标导航。然而,现有方法多使用纯文本查询,在细粒度物体类别上可靠性下降。本文提出IMPRINT框架,通过引入网络获取的图像来增强文本查询,提升语义地图中的物体定位能力。检索图像经视觉-语言模型编码后,与语义地图匹配生成相似性图,并聚合得到上下文感知的定位结果。该方法无需训练或修改底层导航策略。为评估长尾性能,我们构建了基于Habitat合成场景的新基准HSSD-rare,包含语义细粒度子类别。在OVON和HSSD-rare上,图像条件查询均显著提升物体定位精度并带来端到端导航性能提升。进一步分析表明,定位增益能否转化为导航性能提升,关键取决于下游检测质量,揭示了长尾具身导航中的系统瓶颈。

原文摘要 · Abstract (English)

Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (ObjectNav). However, existing approaches typically depend on text-only queries, which become less reliable as semantic specificity increases toward fine-grained object categories. We introduce IMPRINT, a zero-shot plug-and-play framework that enriches textual object queries with web-sourced images to improve grounding in queryable maps. Retrieved images are encoded using a vision-language model, matched against the semantic map to produce similarity maps, and aggregated to yield context-aware localization. Notably, this requires no training or modification of the underlying navigation policy. To explicitly evaluate long-tail behavior, we present HSSD-rare, a new ObjectNav benchmark built on Habitat Synthetic Scenes and featuring semantically specific subcategories. Across both OVON and HSSD-rare, image-conditioned queries consistently improve object grounding and yield end-to-end navigation gains. Further analysis reveals that translating localization gains to navigation performance depends critically on downstream detection quality, highlighting a key systems bottleneck in long-tail embodied navigation.

物体导航视觉语言模型长尾问题零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。