arXiv:2602.18961cs.CVcs.SY2026-02被引 1

融合深度信息与分割模型,提升铁路道砟不足检测的准确性。

Depth-Enhanced YOLO-SAM2 Detection for Reliable Ballast Insufficiency Identification

  • 用深度图校正传感器误差,结合几何分析提升定位精度
  • 召回率从0.49升至0.80,F1分数超0.80
  • 适合铁路安全巡检、视觉模糊场景下的自动化检测

本文提出一种基于RGB-D数据的深度增强型YOLO-SAM2框架,用于铁路轨道道砟不足检测。尽管YOLOv8具备良好定位能力,但仅依赖RGB图像的模型存在安全隐患,其精确率虽达0.99,但召回率仅为0.49,因易误判不足道砟为充足。为提升可靠性,本文引入基于轨枕对齐的深度校正流程,通过多项式建模、RANSAC与时间平滑有效补偿RealSense传感器的空间畸变。结合SAM2分割进一步细化感兴趣区域掩码,实现轨枕与道砟轮廓的精准提取,支持几何分类。在实地采集的俯视RGB-D数据上实验表明,采用深度增强配置后,根据边界框采样方式(AABB或RBB)及几何准则,召回率从0.49提升至最高0.80,F1分数由0.66增至超过0.80。结果证明,将深度校正与YOLO-SAM2融合可显著提升自动化道砟检测的鲁棒性与可靠性,尤其适用于视觉模糊或安全关键场景。

原文摘要 · Abstract (English)

This paper presents a depth-enhanced YOLO-SAM2 framework for detecting ballast insufficiency in railway tracks using RGB-D data. Although YOLOv8 provides reliable localization, the RGB-only model shows limited safety performance, achieving high precision (0.99) but low recall (0.49) due to insufficient ballast, as it tends to over-predict the sufficient class. To improve reliability, we incorporate depth-based geometric analysis enabled by a sleeper-aligned depth-correction pipeline that compensates for RealSense spatial distortion using polynomial modeling, RANSAC, and temporal smoothing. SAM2 segmentation further refines region-of-interest masks, enabling accurate extraction of sleeper and ballast profiles for geometric classification. Experiments on field-collected top-down RGB-D data show that depth-enhanced configurations substantially improve the detection of insufficient ballast. Depending on bounding-box sampling (AABB or RBB) and geometric criteria, recall increases from 0.49 to as high as 0.80, and F1-score improves from 0.66 to over 0.80. These results demonstrate that integrating depth correction with YOLO-SAM2 yields a more robust and reliable approach for automated railway ballast inspection, particularly in visually ambiguous or safety-critical scenarios.

目标检测铁路巡检深度感知分割模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。