让无人机通过自然语言精准导航,提升复杂场景理解能力。
HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones
- 分层跨粒度对比与匹配学习,无需精确分割即可捕捉局部到全局语义。
- 图像检索召回率28.8%,文本检索14.7%,零样本泛化达39.93%均值召回。
- 适用于动态飞行环境,尤其适合语言不完整或模糊的指令理解。
自然语言引导无人机(NLGD)为目标匹配与导航提供了新范式。然而,无人机场景中视野广阔且语义组合复杂,对视觉-语言理解构成挑战。主流视觉-语言模型强调全局对齐,缺乏细粒度语义;现有分层方法依赖精确实体划分和严格包含关系,在动态环境中效果受限。为此,我们提出分层跨粒度对比与匹配学习(HCCM)框架,包含两部分:(1) 区域-全局图像-文本对比学习(RG-ITC),避免精确场景划分,通过局部视觉区域与全局文本之间的双向对比,捕捉层次化局部到全局语义;(2) 区域-全局图像-文本匹配(RG-ITM),摒弃刚性约束,评估全局跨模态表示中局部语义的一致性,增强组合推理能力。此外,无人机文本描述常不完整或模糊,影响对齐稳定性。HCCM引入动量对比与蒸馏(MCD)机制提升鲁棒性。在GeoText-1652数据集上,HCCM实现28.8%(图像检索)和14.7%(文本检索)的Recall@1,达到当前最优水平。在未见的ERA数据集上,零样本泛化表现优异,平均召回率达39.93%,优于微调基线。
原文摘要 · Abstract (English)
Natural Language-Guided Drones (NLGD) provide a novel paradigm for tasks such as target matching and navigation. However, the wide field of view and complex compositional semantics in drone scenarios pose challenges for vision-language understanding. Mainstream Vision-Language Models (VLMs) emphasize global alignment while lacking fine-grained semantics, and existing hierarchical methods depend on precise entity partitioning and strict containment, limiting effectiveness in dynamic environments. To address this, we propose the Hierarchical Cross-Granularity Contrastive and Matching learning (HCCM) framework with two components: (1) Region-Global Image-Text Contrastive Learning (RG-ITC), which avoids precise scene partitioning and captures hierarchical local-to-global semantics by contrasting local visual regions with global text and vice versa; (2) Region-Global Image-Text Matching (RG-ITM), which dispenses with rigid constraints and instead evaluates local semantic consistency within global cross-modal representations, enhancing compositional reasoning. Moreover, drone text descriptions are often incomplete or ambiguous, destabilizing alignment. HCCM introduces a Momentum Contrast and Distillation (MCD) mechanism to improve robustness. Experiments on GeoText-1652 show HCCM achieves state-of-the-art Recall@1 of 28.8% (image retrieval) and 14.7% (text retrieval). On the unseen ERA dataset, HCCM demonstrates strong zero-shot generalization with 39.93% mean recall (mR), outperforming fine-tuned baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。