arXiv:2508.16661cs.CV2025-08被引 7

用视觉语言模型让3D打印质量评估可解释,像人一样给出理由。

QA-VLM: Providing human-interpretable quality assessment for wire-feed laser additive manufacturing parts with Vision Language Models

  • 用视觉语言模型结合专业文献知识,生成可理解的评估理由。
  • 在24个激光送丝沉积样品上验证,解释一致性优于通用模型。
  • 适合需要信任和透明度的工业质检场景,如高端制造。

增材制造中的基于图像的质量评估通常高度依赖熟练操作员的经验和持续关注。尽管机器学习与深度学习方法已被引入以辅助该任务,但它们通常提供黑箱输出且缺乏可解释性,限制了其在实际场景中的可信度和采纳。本文提出一种新型的QA-VLM框架,利用视觉语言模型(VLMs)的注意力机制与推理能力,并融合从同行评审期刊中提炼的应用领域特定知识,生成人类可理解的质量评估结论。在24个通过激光送丝直接能量沉积(DED-LW)工艺制备的单焊道样品上进行评估,本框架在解释质量的有效性和一致性方面均优于现成的VLMs。结果表明,该方法具有实现增材制造中可信、可解释质量评估的潜力。

原文摘要 · Abstract (English)

Image-based quality assessment (QA) in additive manufacturing (AM) often relies heavily on the expertise and constant attention of skilled human operators. While machine learning and deep learning methods have been introduced to assist in this task, they typically provide black-box outputs without interpretable justifications, limiting their trust and adoption in real-world settings. In this work, we introduce a novel QA-VLM framework that leverages the attention mechanisms and reasoning capabilities of vision-language models (VLMs), enriched with application-specific knowledge distilled from peer-reviewed journal articles, to generate human-interpretable quality assessments. Evaluated on 24 single-bead samples produced by laser wire direct energy deposition (DED-LW), our framework demonstrates higher validity and consistency in explanation quality than off-the-shelf VLMs. These results highlight the potential of our approach to enable trustworthy, interpretable quality assessment in AM applications.

质量评估视觉语言模型增材制造可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。