用大模型自动设计机械零件,效果媲美专业软件。
Foundation Models for Automatic CAD Generation

- 融合文本与视觉模型,通过多轮迭代优化生成3D零件
- 最高达98.97%的模型成功生成完整网格,领先模型100%无漏洞
- 适合工业自动化设计,尤其对对称结构有明显短板
大型语言模型(LLMs)和视觉语言模型(VLMs)的进步使得从自然语言描述自动生成参数化3D机械零件成为可能。本文通过统一评估流程和涵盖97个工程设计问题的精选基准,实证研究了基础模型在自动计算机辅助设计(CAD)生成中的表现。提出LLMForge框架,集成JSON-schema校验、分析特征评分、网格合成与多轮迭代优化,分别在两种评审模式下测试:IterTracer使用基于Phong着色的光线追踪渲染器,结合轮廓交并比、孔洞可见性等分析指标提供轻量级几何反馈;IterVision则引入大视觉语言模型Qwen2.5-VL-72B作为语义评议员,通过链式思维视觉推理评估空间一致性和设计意图。在四类典型几何结构(带孔板、多特征盒体、带法兰圆柱、L型支架)上评估七种基础模型。在IterTracer下,前四名模型平均得分稳定在[0.885, 0.890],网格生成成功率高达98.97%,表明微调后的紧凑模型可媲美更大系统。在IterVision中,最优模型实现100%水密网格生成,但对旋转对称结构(如圆柱)存在系统性困难,视觉与语义评分差异显著。论文还讨论了基准设计、失败模式、提示工程及对工业流程的启示。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs from natural-language specifications. This chapter presents an empirical study of foundation models for automatic Computer-Aided Design (CAD) generation of mechanical parts, using a unified evaluation pipeline and a curated benchmark of 97 engineering design problems. We introduce LLMForge, a multi-model text-to-CAD framework integrating JSON-schema validation, analytic feature scoring, mesh synthesis, and multi-round iterative refinement, studied under two critique regimes. IterTracer uses a Phong-shaded ray-trace renderer with analytic visual metrics (silhouette IoU, hole visibility, edge clearance, aspect-ratio conformance) for lightweight geometry-aware feedback across rounds. IterVision replaces the analytic scorer with a VLM semantic critic (Qwen2.5-VL-72B) that evaluates rendered views via chain-of-thought visual reasoning, assessing spatial coherence and design intent. On a benchmark spanning four canonical geometry families (plates with holes and bolt circles, multi-feature boxes, flanged cylinders, and L-brackets), we evaluate seven foundation models: DeepSeek-V3.2, Qwen3-235B-A22B, Llama-3.3-70B, Gemma-3-27B, GLM-4.5, MiniMax-M2.1, and INTELLECT. Under IterTracer, the four highest-ranked models form a tight cluster (overall mean in [0.885, 0.890]) with 98.97% mesh success, showing that compact instruction-tuned models can match substantially larger systems. VLM-based critique in IterVision yields 100% watertight mesh generation on the leading model while surfacing systematic difficulty on rotationally symmetric geometries such as cylinders, where visual and semantic scoring diverge most. We discuss benchmark design, failure modes, CAD-oriented prompting, and implications for industrial workflows and scalable automated mechanical design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。