用大模型引导特征渐进对齐,提升多模态无人机目标检测精度
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
- 利用大语言模型生成细粒度描述并提取语义特征,指导跨模态对齐
- 三阶段渐进对齐:语义对齐、显式空间对齐、隐式空间对齐,显著减少模态差异
- 在两个公开无人机数据集上超越现有方法,适合多模态感知研究者
现有多模态无人机目标检测方法常忽略模态间的语义鸿沟,导致难以实现精准的语义与空间对齐,限制检测性能。为此,本文提出一种大语言模型(LLM)引导的渐进式特征对齐网络LPANet,利用大语言模型提取的语义特征,逐步引导跨模态的语义与空间对齐。首先,通过ChatGPT生成每类物体的细粒度文本描述,并用MPNet提取语义特征;随后设计语义对齐模块(SAM),拉近物体语义与多模态视觉特征距离;其次,设计显式空间对齐模块(ESM),将语义关系融入特征级偏移估计,缓解粗粒度空间错位;最后,设计隐式空间对齐模块(ISM),利用跨模态相关性聚合邻域关键特征,实现隐式空间对齐。在两个公开的多模态无人机目标检测数据集上进行充分实验,结果表明该方法优于当前最优多模态无人机检测器。
原文摘要 · Abstract (English)
Existing multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial alignments, limiting detection performance. To address this problem, we propose a Large Language Model (LLM) guided Progressive feature Alignment Network called LPANet, which leverages the semantic features extracted from a large language model to guide the progressive semantic and spatial alignment between modalities for multimodal UAV object detection. To employ the powerful semantic representation of LLM, we generate the fine-grained text descriptions of each object category by ChatGPT and then extract the semantic features using the large language model MPNet. Based on the semantic features, we guide the semantic and spatial alignments in a progressive manner as follows. First, we design the Semantic Alignment Module (SAM) to pull the semantic features and multimodal visual features of each object closer, alleviating the semantic differences of objects between modalities. Second, we design the Explicit Spatial alignment Module (ESM) by integrating the semantic relations into the estimation of feature-level offsets, alleviating the coarse spatial misalignment between modalities. Finally, we design the Implicit Spatial alignment Module (ISM), which leverages the cross-modal correlations to aggregate key features from neighboring regions to achieve implicit spatial alignment. Comprehensive experiments on two public multimodal UAV object detection datasets demonstrate that our approach outperforms state-of-the-art multimodal UAV object detectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。