用视觉语言模型连接材料微观结构与专家知识,实现无需重训练的缺陷识别。
Linking heterogeneous microstructure informatics with expert characterization knowledge through customized and hybrid vision-language representations for industrial qualification
- 融合视觉与文本特征,构建定制化跨模态表示。
- 零样本分类准确区分合格与缺陷样品,性能优于通用模型。
- 适合材料工程、工业质检等需人机协同决策的场景。
先进材料的快速可靠定级仍是工业制造中的瓶颈,尤其针对非传统增材制造工艺生成的异质结构。本文提出一种新框架,通过定制化混合视觉-语言表征(VLR),将微观结构信息与多种专家表征知识相联结。结合深度语义分割与预训练多模态模型(CLIP和FLAVA),将视觉微观结构数据与文本专家评估编码为共享表征。为克服通用嵌入的局限性,设计了一种基于相似性的定制表征,融合专家标注图像及其文本描述中的正负样本,实现通过净相似度评分进行零样本分类。在增材制造金属基复合材料数据集上的验证表明,该框架可有效区分不同表征标准下的合格与缺陷样本。对比分析显示,FLAVA模型具有更高视觉敏感性,而CLIP模型更一致地匹配文本标准。Z-score标准化根据本地数据分布调整单模态与跨模态相似度分数,提升混合框架中对齐与分类效果。该方法增强了定级流程的可追溯性与可解释性,支持无需任务特定模型重训练的人机协同决策。通过促进原始数据与专家知识间的语义互操作性,为工程信息学中可扩展、领域自适应的定级策略提供支持。
原文摘要 · Abstract (English)
Rapid and reliable qualification of advanced materials remains a bottleneck in industrial manufacturing, particularly for heterogeneous structures produced via non-conventional additive manufacturing processes. This study introduces a novel framework that links microstructure informatics with a range of expert characterization knowledge using customized and hybrid vision-language representations (VLRs). By integrating deep semantic segmentation with pre-trained multi-modal models (CLIP and FLAVA), we encode both visual microstructural data and textual expert assessments into shared representations. To overcome limitations in general-purpose embeddings, we develop a customized similarity-based representation that incorporates both positive and negative references from expert-annotated images and their associated textual descriptions. This allows zero-shot classification of previously unseen microstructures through a net similarity scoring approach. Validation on an additively manufactured metal matrix composite dataset demonstrates the framework's ability to distinguish between acceptable and defective samples across a range of characterization criteria. Comparative analysis reveals that FLAVA model offers higher visual sensitivity, while the CLIP model provides consistent alignment with the textual criteria. Z-score normalization adjusts raw unimodal and cross-modal similarity scores based on their local dataset-driven distributions, enabling more effective alignment and classification in the hybrid vision-language framework. The proposed method enhances traceability and interpretability in qualification pipelines by enabling human-in-the-loop decision-making without task-specific model retraining. By advancing semantic interoperability between raw data and expert knowledge, this work contributes toward scalable and domain-adaptable qualification strategies in engineering informatics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。