用多模态思维链让AI读懂结构损伤并用自然语言描述
SDIGLM: Leveraging Large Language Models and Multi-Modal Chain of Thought for Structural Damage Identification
- 结合语义分割与提示工程,生成视觉和语言双通道推理链条
- 在多种基础设施上实现95.24%的损伤识别准确率
- 适合需要解释性报告的工程检测与智能诊断场景
基于计算机视觉的结构损伤识别模型虽具备较高分类与定位精度,但存在损伤类型识别能力有限、缺乏语言描述能力等瓶颈。本文提出SDIGLM,基于开源VisualGLM-6B架构,引入基于U-Net的语义分割模块生成缺陷分割图作为视觉思维链(CoT),并通过多轮对话微调数据集与提示工程构建语言思维链。该多模态思维链使模型在不同基础设施上达到95.24%的识别准确率,并能有效描述孔洞尺寸、裂纹方向、腐蚀程度等损伤特征。
原文摘要 · Abstract (English)
Existing computer vision(CV)-based structural damage identification models demonstrate notable accuracy in categorizing and localizing damage. However, these models present several critical limitations that hinder their practical application in civil engineering(CE). Primarily, their ability to recognize damage types remains constrained, preventing comprehensive analysis of the highly varied and complex conditions encountered in real-world CE structures. Second, these models lack linguistic capabilities, rendering them unable to articulate structural damage characteristics through natural language descriptions. With the continuous advancement of artificial intelligence(AI), large multi-modal models(LMMs) have emerged as a transformative solution, enabling the unified encoding and alignment of textual and visual data. These models can autonomously generate detailed descriptive narratives of structural damage while demonstrating robust generalization across diverse scenarios and tasks. This study introduces SDIGLM, an innovative LMM for structural damage identification, developed based on the open-source VisualGLM-6B architecture. To address the challenge of adapting LMMs to the intricate and varied operating conditions in CE, this work integrates a U-Net-based semantic segmentation module to generate defect segmentation maps as visual Chain of Thought(CoT). Additionally, a multi-round dialogue fine-tuning dataset is constructed to enhance logical reasoning, complemented by a language CoT formed through prompt engineering. By leveraging this multi-modal CoT, SDIGLM surpasses general-purpose LMMs in structural damage identification, achieving an accuracy of 95.24% across various infrastructure types. Moreover, the model effectively describes damage characteristics such as hole size, crack direction, and corrosion severity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。