用问答模型自动分析无损检测图像,提升效率与准确性。
A Visual Question Answering Model to Automate Nondestructive Evaluation Image Analysis
- 结合ResNet-50与GPT-2,实现图像理解与自然语言回答。
- 支持交互式提问,可精准定位缺陷位置或判断是否存在裂纹。
- 适用于现场检测场景,减少人工误判,提升操作便利性。
本研究提出一种专用于无损检测(NDE)应用的视觉问答(VQA)模型。该模型使检测人员可通过交互式提问,如‘是否存在裂纹’或‘缺陷位于何处’,从模型获取精准答案。系统融合基于ResNet-50的图像特征提取与基于GPT-2的语言生成能力,实现对图像内容的准确理解与信息反馈。通过直接问答交互,该模型显著提升检测效率,降低潜在错误,并增强在实际现场环境中的可用性。
原文摘要 · Abstract (English)
This study introduces a Visual Question Answering model designed specifically for nondestructive evaluation applications. VQA models allow inspectors to interactively query NDE images, asking targeted questions like, Is there a crack or Where is the defect located and receive precise answers from the model. Leveraging deep learning and natural language processing, the developed system integrates image feature extraction (via a ResNet-50 model) and language generation capabilities (via GPT-2) to provide accurate, informative feedback. By enabling direct question-and-answer interactions, this VQA model significantly improves inspection efficiency, reduces potential errors, and enhances usability in practical field scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。