用合成数据和规则引导强化学习,让小模型在晶圆缺陷分析上超越大厂模型。
WaferSAGE: Large Language Model-Powered Wafer Defect Analysis via Synthetic Data Generation and Rubric-Guided Reinforcement Learning

- 三阶段合成流水线生成带结构化规则的缺陷描述与问答对。
- 40亿参数模型达6.493分(LLM-Judge),逼近谷歌Gemini-3-Flash。
- 适合关注工业视觉、隐私安全与低成本部署的芯片制造团队。
我们提出WaferSAGE框架,用于小规模视觉语言模型在晶圆缺陷视觉问答任务中的应用。针对半导体制造中数据稀缺问题,设计了三阶段合成流程:从少量标注晶圆图出发,通过聚类清洗标签噪声,再利用视觉语言模型生成完整缺陷描述,并转化为结构化评估规则;这些规则指导生成涵盖缺陷类型识别、空间分布、形貌特征与根因分析的VQA对。采用基于贝叶斯优化的双评估框架,使规则指标与LLM-Judge评分对齐,实现可靠自动化评估。结合组序列策略优化(GSPO)的课程式强化学习与规则对齐奖励,40亿参数的Qwen3-VL模型取得6.493分(LLM-Judge),接近Gemini-3-Flash的7.149分,同时支持完全本地化部署。结果表明,经过领域特训的小模型可在专业工业视觉理解任务中超越专有大模型,为半导体制造提供可私密、低成本的可行路径。
原文摘要 · Abstract (English)
We present WaferSAGE, a framework for wafer defect visual question answering using small vision-language models. To address data scarcity in semiconductor manufacturing, we propose a three-stage synthesis pipeline incorporating structured rubric generation for precise evaluation. Starting from limited labeled wafer maps, we employ clustering-based cleaning to filter label noise, then generate comprehensive defect descriptions using vision-language models, which are converted into structured evaluation rubrics criteria. These rubrics guide the synthesis of VQA pairs, ensuring coverage across defect type identification, spatial distribution, morphology, and root cause analysis. Our dual assessment framework aligns rule-based metrics with LLM-Judge scores via Bayesian optimization, enabling reliable automated evaluation. Through curriculum-based reinforcement learning with Group Sequence Policy Optimization (GSPO) and rubric-aligned rewards, our 4B-parameter Qwen3-VL model achieves a 6.493 LLM-Judge score, closely approaching Gemini-3-Flash (7.149) while enabling complete on-premise deployment. We demonstrate that small models with domain-specific training can surpass proprietary large models in specialized industrial visual understanding, offering a viable path for privacy-preserving, cost-effective deployment in semiconductor manufacturing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。