评测生成的CAD是否真能用,不只看长得像。
CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

- 分两轨测试参数化建模与装配推理能力
- 复杂编辑和物理仿真匹配仍难达标
- 适合关注CAD生成真实工程可用性的研究者
一个CAD模型仅外观正确并不等于工程可用。它必须满足设计要求、对参数变化有可预测响应、支持可控编辑、在指定分析下结构响应与参考一致,并通过有效接头连接其他部件。我们提出CADEngBench,一个双轨基准测试:CADEngBench-P评估300个参数化零件,涵盖600项任务(零起点建模与功能编辑),通过边界表示有效性、工程与DFM检查、参数族扰动、功能编辑及CalculiX中的匹配线性静态有限元分析(FEA);CADEngBench-A评估150对部件,包含关节检索排序、精确面边对齐、关节坐标系预测与运动学验证。八种多模态代码模型中,修复已有CAD比从头生成更易,但复杂编辑与匹配FEA依然困难。装配预测常定位到区域却无法恢复实际关节或配合实体。结果表明,CAD评估必须检验工程行为而非仅外观。
原文摘要 · Abstract (English)
A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, support controlled edits, match a reference structural response under a declared analysis, and connect to other parts through valid joints. We present CADEngBench, a two-track benchmark for these capabilities. CADEngBench-P evaluates 300 parametric parts, each used for one zero-to-CAD task and one functional-editing task (600 tasks in total), through boundary-representation (B-Rep) validity, engineering and DFM checks, parameter-family perturbations, functional editing, and matched linear-static FEA in CalculiX. CADEngBench-A evaluates 150 body pairs through ranked joint retrieval, exact face-and-edge grounding, joint-frame prediction, and kinematic verification. Across eight multimodal, code-capable models, editing supplied CAD is substantially easier than generating it, while complex edits and matched FEA remain difficult. Assembly predictions often locate the relevant region but fail to recover the recorded joint or mating entities. These results show that CAD evaluation must test engineering behavior rather than appearance alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。