首个评估文本生成参数化CAD的基准,覆盖从基础到真实场景的复杂设计。
Text2CAD-Bench: A Benchmark for LLM-based Text-to-Parametric CAD Generation

- 构建600例多层级案例,包含专家与非专家双风格提示。
- 模型在复杂拓扑和高级功能上性能显著下降,最高误差超40%。
- 适合研究生成式AI与工业设计融合的学者与工程师。
文本到CAD生成旨在通过自然语言创建参数化CAD模型,实现快速原型设计与直观设计流程。然而现有基准仅聚焦基本几何体与简单草图-拉伸序列,缺乏真实应用所需的高级特征,且仅涵盖传统机械部件。本文提出Text2CAD-Bench,首个系统评估文本生成参数化CAD在几何复杂度与应用多样性上的基准。该基准包含600个人工精选案例,分为四个层级:L1-L2涵盖基础几何与标准特征,L3引入复杂拓扑与自由曲面,L4拓展至机械以外的真实世界领域。每例均配备双风格提示:模拟非专家的几何描述,以及符合专家规范的程序化序列。评估主流通用LLM与领域专用模型发现,当前模型在基础几何上表现尚可,但在复杂拓扑与高级功能上性能显著下降。我们已发布该基准以推动文本生成参数化CAD研究进展。
原文摘要 · Abstract (English)
Text-to-CAD generation aims to create parametric CAD models from natural language, enabling rapid prototyping and intuitive design workflows. However, existing benchmarks focus on basic primitives and simple sketch-extrude sequences, lacking advanced features essential for real-world applications and covering only traditional mechanical parts. We introduce Text2CAD-Bench, the first benchmark systematically evaluating text-to-CAD across geometric complexity and application diversity. Our benchmark comprises 600 human-curated examples spanning four levels: L1-L2 cover fundamental geometry with standard features, L3 introduces complex topology and freeform surfaces, and L4 extends to real-world domains beyond mechanical parts. Each example pairs dual-style prompts -- geometric descriptions mimicking non-expert users, and procedural sequences aligned with expert-level conventions. Evaluating mainstream general LLMs and domain-specific models, we find that current models perform reasonably on basic geometry but degrade substantially on complex topology and advanced features. We release our benchmark to drive progress in text-to-CAD research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。