让在线模型学会利用未来帧信息,提升3D目标检测精度。
Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
- 用稀疏查询机制重建未来特征,无需严格对齐帧
- 在nuScenes上提升1.3点mAP和NDS,速度估计最准
- 适合追求高精度的自动驾驶感知系统
基于摄像头的时序3D目标检测在自动驾驶中表现优异,离线模型通过使用未来帧提升了准确率。知识蒸馏(KD)可将离线教师模型中的丰富信息迁移到在线学生模型。然而现有方法忽略未来帧,主要关注严格帧对齐下的空间特征蒸馏或时序关系蒸馏,导致在线模型难以有效学习未来知识。为此,我们提出稀疏查询式未来时序知识蒸馏(FTKD),将离线教师模型的未来帧知识有效迁移至在线学生模型。具体地,设计未来感知特征重建策略,使学生模型在无严格帧对齐条件下捕捉未来特征;进一步引入未来引导的logit蒸馏,利用教师模型稳定的前景与背景上下文。FTKD应用于两个高性能3D目标检测基线,在nuScenes数据集上实现最高1.3 mAP和1.3 NDS提升,且速度估计最精确,推理成本未增加。
原文摘要 · Abstract (English)
Camera-based temporal 3D object detection has shown impressive results in autonomous driving, with offline models improving accuracy by using future frames. Knowledge distillation (KD) can be an appealing framework for transferring rich information from offline models to online models. However, existing KD methods overlook future frames, as they mainly focus on spatial feature distillation under strict frame alignment or on temporal relational distillation, thereby making it challenging for online models to effectively learn future knowledge. To this end, we propose a sparse query-based approach, Future Temporal Knowledge Distillation (FTKD), which effectively transfers future frame knowledge from an offline teacher model to an online student model. Specifically, we present a future-aware feature reconstruction strategy to encourage the student model to capture future features without strict frame alignment. In addition, we further introduce future-guided logit distillation to leverage the teacher's stable foreground and background context. FTKD is applied to two high-performing 3D object detection baselines, achieving up to 1.3 mAP and 1.3 NDS gains on the nuScenes dataset, as well as the most accurate velocity estimation, without increasing inference cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。