arXiv:2410.12143cs.CV2024-10被引 7

无需人工对齐,用多尺度专家网络实现红外可见光视频目标检测

Mixture of Scale Experts for Alignment-free RGBT Video Object Detection and A Unified Benchmark

  • 采用多尺度专家网络自动捕捉可见光与红外图像的尺度差异
  • 动态路由机制根据输入自适应选择最佳专家,提升检测精度
  • 引入可变形卷积缓解空间位置错位问题,适合真实复杂场景

现有可见光-热红外视频目标检测方法高度依赖人工图像对齐,既耗时又难扩展。为此,本文提出混合尺度专家网络(MSENet),通过在不同感知尺度上训练多个专家,无需显式对齐即可捕捉双模态图像间的尺度差异。每个专家能识别并量化图像对中的尺度不一致,动态路由机制则根据输入特征自适应分配权重,选择最优专家。为应对弱对齐位置偏差,网络中嵌入可变形卷积,学习双模态间的位置偏移,降低空间错位影响。此外,本文构建了一个统一基准数据集,涵盖11类常见物体,共60,988张图像和271,835个目标实例,覆盖日常生活与自然环境多种场景,具有高内容多样性和复杂性。

原文摘要 · Abstract (English)

Existing RGB-Thermal Video Object Detection (RGBT VOD) methods predominantly rely on the manual alignment of image pairs, that is both labor-intensive and time-consuming. This dependency significantly restricts the scalability and practical applicability of these methods in real-world scenarios. To address this critical limitation, we propose a novel framework termed the Mixture of Scale Experts Network (MSENet). MSENet integrates multiple experts trained at different perceptual scales, enabling the capture of scale discrepancies between RGB and thermal image pairs without the need for explicit alignment. Specifically, to address the issue of unaligned scales, MSENet introduces a set of experts designed to perceive the correlation between RGBT image pairs across various scales. These experts are capable of identifying and quantifying the scale differences inherent in the image pairs. Subsequently, a dynamic routing mechanism is incorporated to assign adaptive weights to each expert, allowing the network to dynamically select the most appropriate experts based on the specific characteristics of the input data. Furthermore, to address the issue of weakly unaligned positions, we integrate deformable convolution into the network. Deformable convolution is employed to learn position displacements between the RGB and thermal modalities, thereby mitigating the impact of spatial misalignment. To provide a comprehensive evaluation platform for alignment-free RGBT VOD, we introduce a new benchmark dataset. This dataset includes eleven common object categories, with a total of 60,988 images and 271,835 object instances. The dataset encompasses a wide range of scenes from both daily life and natural environments, ensuring high content diversity and complexity.

RGBT检测多模态可变形卷积无对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。