用视觉语言模型自动识别CAD中的制造特征,无需大量数据或规则。
Leveraging Vision-Language Models for Manufacturing Feature Recognition in CAD Designs
- 通过提示工程让大模型理解多视角图纸,实现零样本特征识别。
- Claude-3.5-Sonnet在特征数量和名称识别上准确率达75%,误报率最低。
- 适合制造业数字化转型中需要快速解析复杂设计的工程师使用。
自动特征识别(AFR)对于将设计知识转化为可执行的制造信息至关重要。传统方法依赖预定义几何规则和大规模数据集,往往耗时且泛化能力差。为解决这一问题,本研究探索视觉语言模型(VLMs)在无需大量训练数据或预设规则的情况下,自动识别各类制造特征的能力。采用多视图查询图像、少样本学习、序列推理与思维链等提示工程策略。在新构建的涵盖机加工、增材制造、钣金成形、模具成型与铸造等多种工艺的CAD数据集上进行评估,该数据集包含不同复杂度的设计。共测试五种VLMs:三种闭源模型(GPT-4o、Claude-3.5-Sonnet、Claude-3.0-Opus)和两种开源模型(LLava、MiniCPM),专家标注了真实特征标签。关键指标包括特征数量准确率、特征名称匹配准确率、幻觉率及平均绝对误差(MAE)。结果显示,Claude-3.5-Sonnet在特征数量准确率(74%)和名称匹配准确率(75%)上表现最佳,且MAE最低(3.2);GPT-4o幻觉率最低(8%);而开源模型幻觉率超过30%,准确率低于40%。研究表明,VLMs在多种制造场景下具备自动化识别CAD特征的潜力。
原文摘要 · Abstract (English)
Automatic feature recognition (AFR) is essential for transforming design knowledge into actionable manufacturing information. Traditional AFR methods, which rely on predefined geometric rules and large datasets, are often time-consuming and lack generalizability across various manufacturing features. To address these challenges, this study investigates vision-language models (VLMs) for automating the recognition of a wide range of manufacturing features in CAD designs without the need for extensive training datasets or predefined rules. Instead, prompt engineering techniques, such as multi-view query images, few-shot learning, sequential reasoning, and chain-of-thought, are applied to enable recognition. The approach is evaluated on a newly developed CAD dataset containing designs of varying complexity relevant to machining, additive manufacturing, sheet metal forming, molding, and casting. Five VLMs, including three closed-source models (GPT-4o, Claude-3.5-Sonnet, and Claude-3.0-Opus) and two open-source models (LLava and MiniCPM), are evaluated on this dataset with ground truth features labelled by experts. Key metrics include feature quantity accuracy, feature name matching accuracy, hallucination rate, and mean absolute error (MAE). Results show that Claude-3.5-Sonnet achieves the highest feature quantity accuracy (74%) and name-matching accuracy (75%) with the lowest MAE (3.2), while GPT-4o records the lowest hallucination rate (8%). In contrast, open-source models have higher hallucination rates (>30%) and lower accuracies (<40%). This study demonstrates the potential of VLMs to automate feature recognition in CAD designs within diverse manufacturing scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。