融合结构先验与语言引导,提升多模态目标检测鲁棒性
SLGNet: Synergizing Structural Priors and Language-Guided Modulation for Multimodal Object Detection
- 用结构感知适配器注入跨模态层次结构信息
- 语言引导调制使模型具备环境感知能力,提升动态场景适应性
- 参数量减少87%仍达新SOTA,适合资源受限部署
利用可见光(RGB)与红外(IR)图像进行多模态目标检测对全天候鲁棒感知至关重要。尽管近期基于适配器的方法能高效迁移预训练的RGB基础模型,但常因追求模型效率而牺牲跨模态结构一致性,导致在高对比度或夜间等域差距大的场景中丢失关键结构线索。同时,传统静态融合机制缺乏环境感知能力,难以应对复杂动态场景变化。为此,本文提出SLGNet,一种参数高效的框架,在冻结的视觉变换器(ViT)基础上融合分层结构先验与语言引导调制。设计结构感知适配器从双模态中提取层次化结构表征,并动态注入ViT以补偿其固有的结构退化。进一步提出语言引导调制模块,利用视觉语言模型生成的结构化描述动态重校准视觉特征,赋予模型强环境感知能力。在LLVIP、FLIR、KAIST和DroneVehicle数据集上的大量实验表明,SLGNet达到新SOTA性能。尤其在LLVIP基准上,mAP达66.1,同时可训练参数相比全微调减少约87%。验证了SLGNet在多模态感知中的鲁棒性与高效性。
原文摘要 · Abstract (English)
Multimodal object detection leveraging RGB and Infrared (IR) images is pivotal for robust perception in all-weather scenarios. While recent adapter-based approaches efficiently transfer RGB-pretrained foundation models to this task, they often prioritize model efficiency at the expense of cross-modal structural consistency. Consequently, critical structural cues are frequently lost when significant domain gaps arise, such as in high-contrast or nighttime environments. Moreover, conventional static multimodal fusion mechanisms typically lack environmental awareness, resulting in suboptimal adaptation and constrained detection performance under complex, dynamic scene variations. To address these limitations, we propose SLGNet, a parameter-efficient framework that synergizes hierarchical structural priors and language-guided modulation within a frozen Vision Transformer (ViT)-based foundation model. Specifically, we design a Structure-Aware Adapter to extract hierarchical structural representations from both modalities and dynamically inject them into the ViT to compensate for structural degradation inherent in ViT-based backbones. Furthermore, we propose a Language-Guided Modulation module that exploits VLM-driven structured captions to dynamically recalibrate visual features, thereby endowing the model with robust environmental awareness. Extensive experiments on the LLVIP, FLIR, KAIST, and DroneVehicle datasets demonstrate that SLGNet establishes new state-of-the-art performance. Notably, on the LLVIP benchmark, our method achieves an mAP of 66.1, while reducing trainable parameters by approximately 87% compared to traditional full fine-tuning. This confirms SLGNet as a robust and efficient solution for multimodal perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。