用视觉语言模型解决多光谱行人检测的模态错位问题
Revisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion
- 利用大模型实现可见光与热成像语义对齐
- 在严重错位数据上仍保持高检测精度
- 无需复杂校准,适合实际部署场景
多光谱行人检测在诸多关键应用中至关重要。然而,在真实环境中数据常出现严重模态错位,导致传统基于对齐数据集的方法失效。本文提出一种新框架,专为处理严重错位的多光谱数据设计,无需昂贵复杂的预处理校准。通过利用大规模视觉语言模型(LVLM)实现可见光与热成像之间的跨模态语义对齐,提升两域间信息一致性,从而增强检测准确率。该方法简化了操作流程,显著扩展了多光谱检测技术在实际场景中的可用性。
原文摘要 · Abstract (English)
Multispectral pedestrian detection is a crucial component in various critical applications. However, a significant challenge arises due to the misalignment between these modalities, particularly under real-world conditions where data often appear heavily misaligned. Conventional methods developed on well-aligned or minimally misaligned datasets fail to address these discrepancies adequately. This paper introduces a new framework for multispectral pedestrian detection designed specifically to handle heavily misaligned datasets without the need for costly and complex traditional pre-processing calibration. By leveraging Large-scale Vision-Language Models (LVLM) for cross-modal semantic alignment, our approach seeks to enhance detection accuracy by aligning semantic information across the RGB and thermal domains. This method not only simplifies the operational requirements but also extends the practical usability of multispectral detection technologies in practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。