提出高分辨率事件流跟踪数据集与高效追踪方法
Event Stream-based Visual Object Tracking: HDETrack V2 and A High-Definition Benchmark
- 分层知识蒸馏+时频变换提升模型时序建模能力
- 在低/高分辨率数据集上均实现领先性能
- 适合事件相机视觉跟踪研究者使用
我们提出一种新型分层知识蒸馏策略,融合相似性矩阵、特征表示和响应图的蒸馏,指导学生Transformer网络学习。通过引入时频变换增强模型捕捉帧间时序依赖的能力。针对测试阶段,设计了新的测试时微调策略,使网络能自适应特定目标,提升跟踪性能与灵活性。鉴于现有事件流跟踪数据集多为低分辨率,我们构建了首个大规模高分辨率事件流跟踪数据集EventVOT,包含1141个视频,涵盖行人、车辆、无人机、乒乓球等多种类别。在低分辨率(FE240hz、VisEvent、FELT)及新提出的高分辨率EventVOT数据集上进行充分实验,全面验证了所提方法的有效性。相关基准数据集与源代码已开源至https://github.com/Event-AHU/EventVOT_Benchmark。
原文摘要 · Abstract (English)
We then introduce a novel hierarchical knowledge distillation strategy that incorporates the similarity matrix, feature representation, and response map-based distillation to guide the learning of the student Transformer network. We also enhance the model's ability to capture temporal dependencies by applying the temporal Fourier transform to establish temporal relationships between video frames. We adapt the network model to specific target objects during testing via a newly proposed test-time tuning strategy to achieve high performance and flexibility in target tracking. Recognizing the limitations of existing event-based tracking datasets, which are predominantly low-resolution, we propose EventVOT, the first large-scale high-resolution event-based tracking dataset. It comprises 1141 videos spanning diverse categories such as pedestrians, vehicles, UAVs, ping pong, etc. Extensive experiments on both low-resolution (FE240hz, VisEvent, FELT), and our newly proposed high-resolution EventVOT dataset fully validated the effectiveness of our proposed method. Both the benchmark dataset and source code have been released on https://github.com/Event-AHU/EventVOT_Benchmark
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。