arXiv:2605.07145cond-mat.mtrl-scics.CV2026-05

用专业数据微调视觉语言模型,提升材料断口形貌识别准确率。

Fine-tuning a vision-language model for fracture-surface morphology recognition

  • 基于1.3万张文献挖掘图像,用GPT生成标注并增强数据。
  • 微调后模型精度达0.92,远超基线模型(最高0.78)。
  • 适合需要自动化断口分析的材料科研与显微镜自主工作流。

视觉语言模型(VLM)在科学图像理解中展现出巨大潜力,但通用模型常缺乏材料表征所需的领域特定视觉知识。本文通过精选13,168张开源文献挖掘的断口表面图像,对开源VLM Qwen3-VL-32B-Instruct进行微调。形态学标注由GPT-5.2-Reasoning(high)从图像及原文摘录中生成,并通过人工补充稀有特征图像和旋转增强进一步丰富数据。微调后的专用模型在100张人工标注基准图像上表现优异,精度达0.92,显著优于基线模型(Qwen3-VL-32B-Instruct为0.35,GPT-5.5-Reasoning(high)为0.58,Gemini 3.1 Pro-Reasoning(high)为0.78)。消融实验表明,人工收集罕见特征图像及图像旋转增强均有助于提升对少见断口形貌特征的识别能力。我们还探讨了将微调模型与专有模型结合,以融合断裂特异性视觉精度与更广域多模态推理能力,实现自主断口分析。该研究展示了通过针对性数据收集与微调,使VLM有效识别新特征并支持自主显微镜流程决策的可行性。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have shown strong potential for scientific image understanding, but general-purpose models often lack the domain-specific visual knowledge required for reliable materials characterization. In this work, we fine-tuned an open-source VLM (Qwen3-VL-32B-Instruct) for fracture-surface image analysis using a curated dataset of 13,168 open-source, literature-mined fracture-surface images. Morphology annotations were generated by GPT-5.2-Reasoning (high) from both the images and relevant excerpts of their source papers, and the dataset was further enriched with targeted manual collection and rotation-based augmentation. The resulting specialist model outperforms flagship proprietary multimodal models on a benchmark of 100 manually annotated images. It achieves a precision of 0.92, compared to 0.35 for the base Qwen3-VL-32B-Instruct, 0.58 for GPT-5.5-Reasoning (high), and 0.78 for Gemini 3.1 Pro-Reasoning (high). Dataset ablations show that manual collection of rare-feature images and augmentation via image rotation are both beneficial to improve recognition of less common fracture morphology features. We further discuss integrated use of the fine-tuned model with proprietary models to combine fracture-specific visual accuracy with broader multimodal reasoning for autonomous fractography. Although focused on fracture-surface images, this work demonstrates how VLMs can be adapted through targeted collection and fine-tuning on novel feature images to recognize those features and support downstream decision-making in autonomous microscopy workflows.

视觉语言模型材料表征断口分析微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。