arXiv:2602.01696cs.CVcs.AI2026-02

用视觉与深度信息融合提升小缺陷检测,精度显著优于现有方法。

Cross-Modal Purification and Fusion for Small-Object RGB-D Transmission-Line Defect Detection

  • 先净化后融合:通过代码本过滤噪声,保留关键缺陷特征。
  • 在94.5%为小目标的数据集上,mAP@50达32.2%,小目标AP达12.5%。
  • 轻量版仅4.9M参数,速度228帧/秒,适合无人机实时检测。

输电线路缺陷检测在无人机自动化巡检中仍具挑战,主要因小尺度缺陷占比高、背景复杂且光照变化大。现有基于RGB的检测器在色度对比弱时难以区分几何细微缺陷与相似背景结构。本文提出CMAFNet,一种跨模态对齐与融合网络,通过“净化-融合”范式整合RGB外观与深度几何信息。CMAFNet包含语义重组成模块,利用学习的代码本进行字典式特征净化,抑制模态特异性噪声同时保留缺陷判别信息;以及上下文语义集成框架,采用部分通道注意力捕捉全局空间依赖,增强结构语义推理。净化阶段的位置归一化强制显式重建驱动的跨模态对齐,确保异构特征融合前的统计兼容性。在TLRGBD基准上(94.5%实例为小目标)的大量实验表明,CMAFNet实现32.2% mAP@50和12.5% APs,分别超越最强基线9.8和4.0个百分点。轻量版本仅需4.9M参数,达到228 FPS,性能超越所有YOLO类检测器,且计算成本远低于基于Transformer的方法。

原文摘要 · Abstract (English)

Transmission line defect detection remains challenging for automated UAV inspection due to the dominance of small-scale defects, complex backgrounds, and illumination variations. Existing RGB-based detectors, despite recent progress, struggle to distinguish geometrically subtle defects from visually similar background structures under limited chromatic contrast. This paper proposes CMAFNet, a Cross-Modal Alignment and Fusion Network that integrates RGB appearance and depth geometry through a principled purify-then-fuse paradigm. CMAFNet consists of a Semantic Recomposition Module that performs dictionary-based feature purification via a learned codebook to suppress modality-specific noise while preserving defect-discriminative information, and a Contextual Semantic Integration Framework that captures global spatial dependencies using partial-channel attention to enhance structural semantic reasoning. Position-wise normalization within the purification stage enforces explicit reconstruction-driven cross-modal alignment, ensuring statistical compatibility between heterogeneous features prior to fusion. Extensive experiments on the TLRGBD benchmark, where 94.5% of instances are small objects, demonstrate that CMAFNet achieves 32.2% mAP@50 and 12.5% APs, outperforming the strongest baseline by 9.8 and 4.0 percentage points, respectively. A lightweight variant reaches 24.8% mAP50 at 228 FPS with only 4.9M parameters, surpassing all YOLO-based detectors while matching transformer-based methods at substantially lower computational cost.

缺陷检测多模态融合小目标检测无人机巡检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。