无需训练即可实现隧道缺陷精准定位与工程级报告生成
Training-Free Tunnel Defect Inspection and Engineering Interpretation via Visual Recalibration and Entity Reconstruction

- 通过视觉一致性重校准粗略缺陷提示,提升复杂场景下的定位精度
- 在可见光、GPR和道路缺陷任务上分别达到0.68、0.78、0.72的F1分数
- 输出结构化缺陷实体,支持工程解释与可读报告,适合基建检测场景
隧道检测需支持缺陷定位、测量、严重性分级及工程文档生成。现有无训练基础模型方法通常仅生成粗粒度开放词汇提议,难以在干扰复杂的隧道场景中直接使用。本文提出无训练框架TunnelMIND:语言引导的缺陷提议不作为最终输出,而是通过密集视觉一致性在推理时进行空间重校准,使粗略语义锚点转化为更可靠的提示,应对隧道特有强负样本。生成的掩码进一步重构为包含类别、位置、几何、严重性和上下文属性的结构化缺陷实体,并在专家知识约束下映射至检索增强解释与工程可读报告生成。在可见光、GPR和道路缺陷任务上,TunnelMIND分别取得0.68、0.78和0.72的F1分数。结果表明,无训练隧道检测可从粗略定位迈向支持工程评估的结构化缺陷证据。
原文摘要 · Abstract (English)
Tunnel inspection requires outputs that can support defect localization, measurement, severity grading, and engineering documentation. Existing training-free foundation-model pipelines usually stop at coarse open-vocabulary proposals, which are difficult to use directly in interference-heavy tunnel scenes. We propose a training-free framework TunnelMIND. Specifically, language-guided defect proposals are not treated as final outputs; instead, their spatial support is recalibrated at inference time through dense visual consistency, so that coarse semantic anchors can be transformed into more reliable prompts under tunnel-specific hard negatives. The resulting masks are further reconstructed into structured defect entities with category, location, geometry, severity, and context attributes, which are then mapped to retrieval-grounded explanation and engineering-readable report generation under expert knowledge constraints. On visible, GPR, and road defect tasks, TunnelMIND achieves F1 scores of 0.68, 0.78, and 0.72, respectively. Overall, TunnelMIND shows that training-free tunnel inspection can move beyond coarse localization toward structured defect evidence for engineering assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。