用大模型+文献检索实现激光3D打印缺陷零样本检测与解释。
Multimodal RAG-driven Anomaly Detection and Classification in Laser Powder Bed Fusion using Large Language Models
- 基于文献图文检索增强生成,无需训练即可识别缺陷
- 在4个实验数据集上准确率提升12%,优于基线模型
- 适合快速响应新工艺的3D打印质量分析场景
增材制造可高效生产复杂结构并减少材料浪费,但存在缺陷与工艺异常问题。本文提出一种多模态检索增强生成框架,通过从科学文献中检索文本与图像信息,实现无需训练的零样本缺陷检测、分类与解释生成。该框架在橡树岭国家实验室提供的4个激光粉末床熔融(L-PBF)数据集上验证,涵盖不同品牌、型号打印机与材料。结果表明其具备跨域适应性与泛化能力。以Qwen2-VL-2B和GPT-4o-mini作为多模态大模型进行对比,GPT-4o-mini在缺陷分类任务中表现更优,且高于比例随机基线。RAG机制使平均准确率提升12%,有效降低幻觉风险并提供额外背景信息。框架支持持续更新,可随新研究成果动态演进。该可扩展、自动化、零样本的方案显著提升增材制造缺陷分析效率与准确性。
原文摘要 · Abstract (English)
Additive manufacturing enables the fabrication of complex designs while minimizing waste, but faces challenges related to defects and process anomalies. This study presents a novel multimodal Retrieval-Augmented Generation-based framework that automates anomaly detection across various Additive Manufacturing processes leveraging retrieved information from literature, including images and descriptive text, rather than training datasets. This framework integrates text and image retrieval from scientific literature and multimodal generation models to perform zero-shot anomaly identification, classification, and explanation generation in a Laser Powder Bed Fusion setting. The proposed framework is evaluated on four L-PBF manufacturing datasets from Oak Ridge National Laboratory, featuring various printer makes, models, and materials. This evaluation demonstrates the framework's adaptability and generalizability across diverse images without requiring additional training. Comparative analysis using Qwen2-VL-2B and GPT-4o-mini as MLLM within the proposed framework highlights that GPT-4o-mini outperforms Qwen2-VL-2B and proportional random baseline in manufacturing anomalies classification. Additionally, the evaluation of the RAG system confirms that incorporating retrieval mechanisms improves average accuracy by 12% by reducing the risk of hallucination and providing additional information. The proposed framework can be continuously updated by integrating emerging research, allowing seamless adaptation to the evolving landscape of AM technologies. This scalable, automated, and zero-shot-capable framework streamlines AM anomaly analysis, enhancing efficiency and accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。