arXiv:2509.06333cs.CVcs.RO2025-09被引 1

融合可见光与热成像,提升复杂环境下行人等脆弱道路使用者的检测准确率。

Multi-Modal Camera-Based Detection of Vulnerable Road Users

  • 采用RGB与热成像双模态输入,微调YOLOv8模型进行检测
  • 热成像模型精度最高,特定数据增强使稀有类召回率提升
  • 适合自动驾驶和智能交通系统中高可靠性目标检测场景

行人、骑行者和摩托车手等脆弱道路使用者(VRUs)占全球交通事故死亡人数的一半以上,但在光照不足、恶劣天气及数据不平衡条件下检测仍具挑战。本文提出一种多模态检测框架,结合可见光(RGB)与热红外成像,并基于预训练的YOLOv8模型进行优化。训练使用KITTI、BDD100K和Teledyne FLIR数据集,通过类别重加权与轻量级数据增强提升少数类表现与模型鲁棒性。实验表明,在640像素分辨率下并部分冻结主干网络可实现精度与效率的最佳平衡;类别加权损失显著提升稀有类别的召回率。结果表明,热成像模型具备最高精度,且从RGB到热成像的增强策略有效提升召回性能,验证了多模态检测在交叉路口提升VRU安全性的潜力。

原文摘要 · Abstract (English)

Vulnerable road users (VRUs) such as pedestrians, cyclists, and motorcyclists represent more than half of global traffic deaths, yet their detection remains challenging in poor lighting, adverse weather, and unbalanced data sets. This paper presents a multimodal detection framework that integrates RGB and thermal infrared imaging with a fine-tuned YOLOv8 model. Training leveraged KITTI, BDD100K, and Teledyne FLIR datasets, with class re-weighting and light augmentations to improve minority-class performance and robustness, experiments show that 640-pixel resolution and partial backbone freezing optimise accuracy and efficiency, while class-weighted losses enhance recall for rare VRUs. Results highlight that thermal models achieve the highest precision, and RGB-to-thermal augmentation boosts recall, demonstrating the potential of multimodal detection to improve VRU safety at intersections.

多模态检测目标检测自动驾驶热成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。