arXiv:2608.09270cs.CVcs.AI2026-08中稿 · the 34th ACM Inter…

解决无人机视角下细粒度跨模态理解的难题,提升模型对细节的识别能力。

GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views

论文配图:GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views
图 1 · 摘自论文原文
  • 引入区域聚焦对齐,强化物体中心对齐并抑制背景干扰
  • 通过语义扰动匹配与纯净原型码本,区分外观相似的细微差异
  • 适用于无人机图像文本检索,尤其在复杂背景中表现优异

无人机视角下的细粒度跨模态理解对空中视觉语言导航至关重要。然而,其广角视野和俯视视角带来双重挑战:宏观上,视觉表征中大量背景杂波导致跨模态注意力错位,模型偏向全局环境相似性而非具体物体细节;微观上,视觉同构性造成歧义,候选对象几何结构相似但属性细微不同。为此,我们提出颗粒度感知的区域对齐与语义原型学习框架(GRASP),通过两项协同策略增强判别能力。首先,引入区域聚焦对齐(RFA)实现以物体为中心的跨模态对齐,同时抑制背景干扰。其次,为应对视觉同构性,提出语义扰动增强匹配(SPEM),利用前景净化的语义原型码本(SPC)构建语义扰动负样本,实现细粒度语义区分。在GeoText-1652基准和未见的ERA数据集上的大量实验表明,GRASP在无人机视角细粒度图文检索任务中达到领先性能,验证了其在空域跨模态理解中的有效性。代码已开源:https://github.com/UCAS-JC/GRASP。

原文摘要 · Abstract (English)

Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and overhead perspective of drone scenarios impose dual challenges on vision-language understanding. At the macro level, overwhelming background clutter in visual representations leads to Cross-Modal Focus Misalignment, where the model prioritizes global environmental similarities over specific object details. At the micro level, Visual Isomorphism creates ambiguity, where candidates share similar geometric structures yet differ only in subtle attributes. To address these challenges, we propose the Granularity-Aware Region Alignment and Semantic Prototype (GRASP) learning framework, enhancing discriminative capability through two synergistic strategies. Specifically, we introduce Region-Focused Alignment (RFA) to promote object-centric cross-modal alignment while suppressing background interference. Concurrently, to tackle visual isomorphism, we propose Semantic Perturbation Enhanced Matching (SPEM), which leverages a foreground-purified Semantic Prototype Codebook (SPC) to construct semantically perturbed negatives for fine-grained semantic discrimination. Extensive experiments on the GeoText-1652 benchmark and the unseen ERA dataset demonstrate that GRASP achieves competitive performance in drone-view fine-grained image-text retrieval, validating its effectiveness for cross-modal understanding in aerial scenarios. Our code implementation is available at https://github.com/UCAS-JC/GRASP.

细粒度识别无人机视觉跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。