提出双相关性超图网络,解决红外可见光视频目标检测中的错位问题。
Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark

- 通过局部区域级对齐和双超图融合,建模时序与跨模态相关性。
- 在VT-VOD50和自建的DVT-VOD1000数据集上达到领先精度。
- 构建含1000段视频、10万+图像对的大规模多场景基准数据集。
RGB-热成像(RGBT)视频目标检测因能克服常规基于RGB的检测在复杂条件下的局限而受到关注。然而,RGBT图像对间常存在空间错位。为此,我们提出双相关性超图网络(DHNet),通过显式建模两种相关性——连续帧间的时序相关性和跨模态特征的空间相关性——捕捉高维互补信息。首先设计基于块的局部对齐模块(PSAM),在局部区域层面逐步对齐多模态特征;随后引入双超图融合模块(DHFM),分别构建时序与跨模态超图,通过双相关性学习提升目标判别力。此外,当前领域缺乏大规模、场景多样的基准数据集。为此,我们构建了DVT-VOD1000,一个包含1,000段视频序列、103,464对RGBT图像的大规模数据集,覆盖校园、公园、交通、农村、夜间、雨雪等多种场景。在VT-VOD50与DVT-VOD1000上的综合实验表明,DHNet实现最优检测性能。数据集与源代码将公开于https://github.com/tzz-ahu/,以支持学术研究。
原文摘要 · Abstract (English)
RGB-Thermal (RGBT) Video Object Detection (VOD) has gained significant traction due to its ability to overcome the limitations of conventional RGB-based VOD under challenging conditions. However, spatial misalignment commonly exists between RGBT image pairs. To address this, we propose a Dual-Correlation Hypergraph Network (DHNet) that captures high-dimensional complementary information by explicitly modeling two types of correlations: temporal correlation across consecutive frames and spatial correlation from cross-modal features. Specifically, we first design a Patch-based Spatial Alignment Module (PSAM) to sequentially align the multimodal features at the local region level. Subsequently, we introduce a Dual Hypergraph Fusion Module (DHFM), which constructs separate temporal and multimodal hypergraphs to enhance object discriminability through dual-correlation learning. Furthermore, the field currently lacks a large-scale, scene-diverse benchmark dataset for comprehensive evaluation. To address this gap, we construct DVT-VOD1000, a large-scale RGBT VOD dataset containing 1,000 video sequences with 103,464 RGBT image pairs. The dataset covers diverse scenarios, including campuses, parks, transportation, rural areas, night scenes, rain, and snow. Comprehensive experiments on VT-VOD50 and our DVT-VOD1000 demonstrate that DHNet achieves state-of-the-art detection accuracy. The dataset and source code will be made publicly available on https://github.com/tzz-ahu/ to support academic research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。