首个大规模管道磁泄漏检测数据集,助力智能巡检算法研发
PipeMFL-240K: A Large-scale Dataset and Benchmark for Object Detection in Pipeline Magnetic Flux Leakage Imaging

- 构建24万张伪彩色图像与20万标注框,覆盖12类缺陷
- 针对极长尾分布与微小目标难题,建立真实复杂场景基准
- 适合工业缺陷检测、智能运维与深度学习研究者使用
管道完整性对工业安全和环境保护至关重要,磁通泄漏(MFL)检测是主要的无损检测技术。尽管深度学习有望实现MFL分析自动化,但缺乏大规模公开数据集和基准测试,导致模型比较和可复现评估困难。本文提出 extbf{PipeMFL-240K},一个大规模、精细标注的数据集与基准,用于复杂场景下的管道MFL伪彩色图像目标检测。该数据集反映真实检测复杂性,面临三大挑战:(i) 覆盖 extbf{12}类缺陷的极长尾分布,(ii) 微小目标占比高,常仅数像素大小,(iii) 类内差异显著。数据集包含 extbf{249,320}张图像与 extbf{200,020}个高质量边界框标注,来自12条总长约 extbf{1,530}公里的管道。通过在先进目标检测器上开展大量实验,建立了基线性能。结果表明,现有模型仍难以应对MFL数据内在特性,仍有巨大提升空间。作为首个同规模、同范围的公开数据集与基准,它为高效管道诊断、维护规划提供基础,并有望推动基于MFL的管道完整性评估算法创新与可复现研究。
原文摘要 · Abstract (English)
Pipeline integrity is critical to industrial safety and environmental protection, with Magnetic Flux Leakage (MFL) detection being a primary non-destructive testing technology. Despite the promise of deep learning for automating MFL interpretation, progress toward reliable models has been constrained by the absence of a large-scale public dataset and benchmark, making fair comparison and reproducible evaluation difficult. We introduce \textbf{PipeMFL-240K}, a large-scale, meticulously annotated dataset and benchmark for complex object detection in pipeline MFL pseudo-color images. PipeMFL-240K reflects real-world inspection complexity and poses several unique challenges: (i) an extremely long-tailed distribution over \textbf{12} categories, (ii) a high prevalence of tiny objects that often comprise only a handful of pixels and (iii) substantial intra-class variability. The dataset contains \textbf{249,320} images and \textbf{200,020} high-quality bounding-box annotations, collected from 12 pipelines spanning approximately \textbf{1,530} km. Extensive experiments are conducted with state-of-the-art object detectors to establish baselines. Results show that modern detectors still struggle with the intrinsic properties of MFL data, highlighting considerable headroom for improvement, while PipeMFL-240K provides a reliable and challenging testbed to drive future research. As the first public dataset and the first benchmark of this scale and scope for pipeline MFL inspection, it provides a critical foundation for efficient pipeline diagnostics as well as maintenance planning and is expected to accelerate algorithmic innovation and reproducible research in MFL-based pipeline integrity assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。