arXiv:2511.05551cs.CV2025-11

用少量样本让视觉语言模型自动判断金属3D打印质量,又准又可解释。

In-Context-Learning-Assisted Quality Assessment Vision-Language Models for Metal Additive Manufacturing

  • 通过上下文学习注入少量示范样本,让模型快速适应特定制造任务。
  • 仅需极少样本即达到传统模型的分类准确率,最高接近95%。
  • 生成人类可读推理过程,适合需要透明决策的工业质检场景。

基于视觉的质量评估在增材制造中通常需要专用机器学习模型和特定数据集,但数据收集与模型训练成本高、耗时长。本文利用视觉语言模型(VLM)的推理能力,结合上下文学习(ICL)提供应用特定知识与示范样本,无需大规模专用数据集即可完成模型训练。我们测试了多种ICL采样策略,以寻找在有限样本下的最优配置。实验在两个VLM(Gemini-2.5-flash 和 Gemma3:27b)上进行,针对线激光直接能量沉积工艺的质量评估任务。结果表明,经ICL辅助的VLM在仅使用极少量样本的情况下,能达到与传统机器学习模型相当的分类准确率,最高达94.8%。此外,相比缺乏可解释性的传统分类模型,VLM能生成人类可理解的推理过程,提升可信度。由于制造业中尚无评估可解释性的标准指标,我们提出“知识相关性”和“推理有效性”两项新指标来衡量支持性理由的质量。结果表明,该方法能在数据稀缺条件下实现高精度且具备有效解释能力的质检。

原文摘要 · Abstract (English)

Vision-based quality assessment in additive manufacturing often requires dedicated machine learning models and application-specific datasets. However, data collection and model training can be expensive and time-consuming. In this paper, we leverage vision-language models' (VLMs') reasoning capabilities to assess the quality of printed parts and introduce in-context learning (ICL) to provide VLMs with necessary application-specific knowledge and demonstration samples. This method eliminates the requirement for large application-specific datasets for training models. We explored different sampling strategies for ICL to search for the optimal configuration that makes use of limited samples. We evaluated these strategies on two VLMs, Gemini-2.5-flash and Gemma3:27b, with quality assessment tasks in wire-laser direct energy deposition processes. The results show that ICL-assisted VLMs can reach quality classification accuracies similar to those of traditional machine learning models while requiring only a minimal number of samples. In addition, unlike traditional classification models that lack transparency, VLMs can generate human-interpretable rationales to enhance trust. Since there are no metrics to evaluate their interpretability in manufacturing applications, we propose two metrics, knowledge relevance and rationale validity, to evaluate the quality of VLMs' supporting rationales. Our results show that ICL-assisted VLMs can address application-specific tasks with limited data, achieving relatively high accuracy while also providing valid supporting rationales for improved decision transparency.

质量评估视觉语言模型上下文学习工业质检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。