arXiv:2608.09522cs.CV2026-08

用三视图融合提升雷达探测软土空洞的精度

TriView-YOLO: Early Multi-View Fusion for Ground Penetrating Radar Cavity Detection in Soft, High-Water-Content Soils

  • 三视角输入通过融合层早期整合,增强弱信号空洞特征
  • 在曼谷软黏土实测数据上达到mAP50 0.558,推理仅3.1毫秒
  • 专为高含水软土设计,适合城市道路地下空洞筛查

在软且高含水土壤中,地表雷达(GPR)信号易被衰减,导致地下空洞反射信号微弱,检测极为困难,但此类地质条件恰恰最易形成空洞。本文提出TriView-YOLO,一种基于多视图的YOLOv12检测器,用于此类土壤中的道路空洞筛查。采用纵向B-scan、水平C-scan和横截面B-scan三个共配准视图,组成9通道输入,通过替换YOLOv12主干的TripleInputConv层实现早期融合;网络其余部分保持不变,仅需在纵向视图上输出边界框。训练使用1600个专家标注的实地样本,主要来自泰国曼谷的城市道路调查,采用车载多通道三维GPR移动测绘系统采集,同时加入日本较硬基底的调查数据用于训练与验证,测试集仅包含曼谷实测数据,土壤为软海洋黏土,含水率80%-140%,地下水位1-2米深。该条件下尚无专用深度学习检测评估报道。在未进行数据增强的纯实地测试集上,模型在三次随机种子下平均mAP50为0.558±0.028,计算量23.6 GFLOPs,单图推理时间3.1毫秒。消融实验表明,移除辅助视图会降低mAP50和召回率;而使用公开或合成图像、DINOv3特征、更大模型规模及COCO预训练均未带来性能提升。

原文摘要 · Abstract (English)

Automated detection of subsurface cavities from Ground Penetrating Radar (GPR) is most difficult in soft, high-water-content ground, where conductive, water-saturated soil attenuates the signal and degrades cavity reflections, yet this is also the condition under which cavities most readily form. This paper proposes TriView-YOLO, a multi-view YOLOv12 detector for road cavity screening in such ground. Three co-registered views (longitudinal B-scan, horizontal C-scan, and cross-section B-scan) form a 9-channel input fused by a TripleInputConv layer that replaces the YOLOv12 stem; the rest of the network is unchanged, and bounding boxes are required on the longitudinal view only. Training used 1,600 expert-verified field samples, principally metropolitan road surveys of Bangkok, Thailand, acquired with a vehicle-mounted multichannel three-dimensional GPR mobile mapping system, with surveys over the firmer subgrades of Japan added to training and validation only. The test set comes exclusively from the Bangkok surveys, over soft marine clay with 80-140% water content and a water table at 1-2 m depth, a ground condition for which no dedicated deep learning cavity-detection evaluation has been reported. On this unaugmented, field-only test set, split randomly within surveys, the proposed model attains mAP50 of 0.558 +/- 0.028 over three seeds at 23.6 GFLOPs and 3.1 ms per image. Ablations show that removing the auxiliary views lowers mAP50 and recall, whereas public and synthetic training images, DINOv3 features, larger model scale, and COCO pretraining bring no gain.

雷达检测空洞识别多视图融合城市道路

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。