arXiv:2506.00365cs.CVeess.SP2025-06

融合可见光与红外图像,用知识蒸馏提升检测精度与速度。

Feature Fusion and Knowledge-Distilled Multi-Modal Multi-Target Detection

  • 双模态输入融合RGB与热成像,提升目标感知能力。
  • 学生模型达教师模型95%精度,推理速度降低50%。
  • 适合嵌入式设备部署,兼顾精度与实时性。

在监控与防御领域,多目标检测与分类(MTD)因异构数据源输入及资源受限嵌入式设备上的算法计算复杂度而面临挑战,尤其对基于AI的解决方案而言。为此,我们提出一种基于特征融合与知识蒸馏的多模态多目标检测框架,利用数据融合提升精度,并通过知识蒸馏优化领域适应性。具体地,该方法在新型融合式多模态模型中同时使用RGB与热成像输入,并结合蒸馏训练流程。我们将问题建模为后验概率优化任务,通过多阶段训练流水线与复合损失函数求解,该损失函数能有效将知识从教师模型迁移至学生模型。实验表明,学生模型达到教师模型约95%的mAP,同时推理时间减少约50%,充分验证其在实际MTD部署场景中的适用性。

原文摘要 · Abstract (English)

In the surveillance and defense domain, multi-target detection and classification (MTD) is considered essential yet challenging due to heterogeneous inputs from diverse data sources and the computational complexity of algorithms designed for resource-constrained embedded devices, particularly for Al-based solutions. To address these challenges, we propose a feature fusion and knowledge-distilled framework for multi-modal MTD that leverages data fusion to enhance accuracy and employs knowledge distillation for improved domain adaptation. Specifically, our approach utilizes both RGB and thermal image inputs within a novel fusion-based multi-modal model, coupled with a distillation training pipeline. We formulate the problem as a posterior probability optimization task, which is solved through a multi-stage training pipeline supported by a composite loss function. This loss function effectively transfers knowledge from a teacher model to a student model. Experimental results demonstrate that our student model achieves approximately 95% of the teacher model's mean Average Precision while reducing inference time by approximately 50%, underscoring its suitability for practical MTD deployment scenarios.

多模态检测知识蒸馏嵌入式部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。