arXiv:2603.17390cs.CV2026-03被引 3

用大模型生成材料图像并自动标注,提升小样本下的分类准确率

FMMC: Harnessing the Power of Foundation Models for Accurate Material Classification

  • 用文本提示融合物体与材质信息,自动生成带标签的材料图像数据
  • 结合大模型先验知识联合微调,使模型在少样本下仍保持高精度
  • 适合做材料识别、数字内容生成的研究者和工程师使用

材料分类在计算机视觉与图形学中日益重要,支撑数字与现实世界应用中的材料属性精准分配。传统上被视为图像分类任务,但因标注数据稀缺,模型精度与泛化能力受限。近期视觉语言基础模型(VLMs)提供了新路径,但现有方法在材料识别上仍表现不佳。本文提出FMMC框架,有效利用基础模型克服数据局限。核心创新包括:(a) 构建鲁棒的图像生成与自动标注流水线,生成以材料为中心的多样化高质量图像,并通过融合物体语义与材质属性的文本提示实现自动标签;(b) 引入先验知识提取策略,将VLM信息蒸馏为先验,并与预训练视觉模型联合微调,兼顾广泛泛化性与材料特异性特征适配。大量实验表明,合成数据有效捕捉真实材料特性,结合VLM先验显著提升最终性能。源代码与数据集将公开。

原文摘要 · Abstract (English)

Material classification has emerged as a critical task in computer vision and graphics, supporting the assignment of accurate material properties to a wide range of digital and real-world applications. While traditionally framed as an image classification task, this domain faces significant challenges due to the scarcity of annotated data, limiting the accuracy and generalizability of trained models. Recent advances in vision-language foundation models (VLMs) offer promising avenues to address these issues, yet existing solutions leveraging these models still exhibit unsatisfying results in material recognition tasks. In this work, we propose a novel framework that effectively harnesses foundation models to overcome data limitations and enhance classification accuracy. Our method integrates two key innovations: (a) a robust image generation and auto-labeling pipeline that creates a diverse and high-quality training dataset with material-centric images, and automatically assigns labels by fusing object semantics and material attributes in text prompts; (b) a prior incorporation strategy to distill information from VLMs, combined with a joint fine-tuning method that optimizes a pre-trained vision foundation model alongside VLM-derived priors, preserving broad generalizability while adapting to material-specific features. Extensive experiments demonstrate significant improvements on multiple datasets. We show that our synthetic dataset effectively captures the characteristics of real world materials, and the integration of priors from vision-language models significantly enhances the final performance. The source code and dataset will be released.

材料分类视觉语言模型数据生成少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。