通过多域多阶跨模态知识蒸馏,提升事件相机目标检测性能。
M^2C-EvDet: Multi-Domain Multi-Order Cross-Modal Knowledge Distillation for Event-based Object Detection

- 基于频域解耦与超图计算,设计双模块蒸馏框架。
- 在DVS128-Gesture和ECP-Object datasets上显著提升检测精度。
- 适合事件相机、低延迟视觉系统研究者参考。
事件相机目标检测(EvDet)作为类生物视觉感知范式,在高时间分辨率和宽动态范围场景中表现优异。然而,事件数据固有的稀疏性与视觉语义不足导致其性能远低于基于帧的检测方法。现有工作虽尝试通过知识蒸馏缓解跨模态差异,但仅关注空间语义或成对关系,难以应对复杂场景。为此,本文提出M^2C-EvDet框架,一种多域多阶跨模态知识蒸馏方法。该框架基于频率学习与超图计算,引入两个专用模块:自适应频域解耦特征蒸馏(AF^2D^2)与多阶关系蒸馏(MORD),有效捕捉事件数据中的多层次语义信息,在DVS128-Gesture与ECP-Object数据集上实现性能突破。
原文摘要 · Abstract (English)
Event-based object Detection (EvDet), as a biologically inspired visual perception paradigm, demonstrates superior performance in scenarios demanding high temporal resolution and a wide dynamic range. Nevertheless, the inherent sparse representations and inadequate visual semantics of event data result in a considerable performance disparity between EvDet and frame-based object detection. Previous works attempt to alleviate this cross-modal discrepancy through knowledge distillation, yet they only focus on spatial visual semantics or pair-wise relational information, thus limiting performance in more complex scenarios. To address this challenge, this paper proposes M^2C-EvDet, a Multi-domain and Multi-order Cross-modal knowledge distillation framework for EvDet. Built upon frequency learning and hypergraph computation, M^2C-EvDet integrates two specialized modules: Adaptive Frequency-Decoupled Feature Distillation (AF^2D^2) and Multi-Order Relational Distillation (MORD).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。