arXiv:2509.06987cs.CVcs.AI2025-09被引 1

融合音视频信息提升铁路缺陷检测精度,减少误报。

FusWay: Multimodal hybrid fusion approach. Application to Railway Defect Detection

  • 结合YOLOv8n与视觉变换器,通过多层特征融合音视频数据。
  • 在真实铁路数据集上,精度和整体准确率提升0.2个百分点。
  • 适合需要高可靠性缺陷检测的工业场景使用。

多模态融合是一种在图像伴随信号或音频的任务中日益流行的技术。尽管基于图像的检测方法如YOLO系列可高效用于缺陷检测,但单一模态方法仍存在局限:当缺陷外观与正常结构元素相似时易产生过检。本文提出一种新型多模态融合架构,基于领域规则,采用YOLOv8n进行快速目标检测,并结合视觉变压器(ViT)提取多层(第7、16、19层)特征图与合成音频表示,实现音频与图像的融合。该方法针对两类缺陷——轨道断裂与表面缺陷。在真实铁路数据集上的实验表明,相比仅使用视觉的方法,该多模态融合方案将精度和整体准确率提升了0.2个百分点。学生未配对t检验确认了平均准确率差异具有统计显著性。

原文摘要 · Abstract (English)

Multimodal fusion is a multimedia technique that has become popular in the wide range of tasks where image information is accompanied by a signal/audio. The latter may not convey highly semantic information, such as speech or music, but some measures such as audio signal recorded by mics in the goal to detect rail structure elements or defects. While classical detection approaches such as You Only Look Once (YOLO) family detectors can be efficiently deployed for defect detection on the image modality, the single modality approaches remain limited. They yield an overdetection in case of the appearance similar to normal structural elements. The paper proposes a new multimodal fusion architecture built on the basis of domain rules with YOLO and Vision transformer backbones. It integrates YOLOv8n for rapid object detection with a Vision Transformer (ViT) to combine feature maps extracted from multiple layers (7, 16, and 19) and synthesised audio representations for two defect classes: rail Rupture and Surface defect. Fusion is performed between audio and image. Experimental evaluation on a real-world railway dataset demonstrates that our multimodal fusion improves precision and overall accuracy by 0.2 points compared to the vision-only approach. Student's unpaired t-test also confirms statistical significance of differences in the mean accuracy.

多模态融合铁路检测视觉变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。