构建真实工业设计意图的参数化CAD建模基准,揭示大模型在实际应用中的局限性。
RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

- 基于真实工业场景设计12,632个任务,支持文本、工程图、产品图等多模态输入。
- 九个前沿大模型在执行率、几何重叠度和视觉语义一致性上无一全优,最高综合得分非单项冠军。
- 发现模型常缺失细节、丢失零件身份或装配错误,说明仅靠可执行性无法评估真实建模能力。
参数化计算机辅助设计(CAD)建模难以用单一指标评估。现有基准多聚焦于合成或纯CAD环境,输入模态有限,且仅关注可执行性与交并比(IoU)。本文提出RealCADBench,一个基于真实工业设计意图的参数化CAD建模基准。该基准包含来自19类工厂自动化任务的12,632个任务,涵盖文本描述、2D工程图、实物产品图及渲染图像,支持零件与装配体建模。我们报告了1,770个任务的评估结果:1,745个零件任务分布在四种输入模式下,以及用于所有装配比较的RCB-Assm25(25个装配任务)。各方法生成FreeCAD API Python代码,由统一运行时执行并导出3D模型。评估指标包括可执行性、实体交并比(Solid IoU)、表面交并比(Surface IoU)和基于评分的视觉-语义身份判断(Judge)。在九个独立前沿大模型中,无一在四项指标上全面领先。六种规模的前沿模型中,可执行性为0.565–0.812,Solid IoU为0.2841–0.5379,Surface IoU为0.112–0.217。综合得分最高的模型与单项指标领导者不同。在RCB-Assm25上,Codex搭配GPT-5.5相比独立GPT-5.5提升了可执行性与两个IoU,但视觉-语义判断分数下降6.98个百分点,表明独立GPT-5.5仍是判断力最强。我们还观察到重复失败模式:细部缺失、零件身份丢失、装配位置错误。结果表明,仅靠可执行性不足以刻画真实CAD建模性能,前沿模型与智能体在可执行性、几何重叠度和视觉-语义一致性方面存在显著差异。
原文摘要 · Abstract (English)
Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD benchmarks often emphasize synthetic or CAD-native settings, limited input modalities, or executability and IoUs alone. We introduce RealCADBench, a benchmark for intent-to-program CAD modeling from real industrial design intents. It contains 12,632 tasks from 19 factory-automation categories and spans text descriptions, 2D engineering drawings, real product pictures, and rendered images for both Part and Assembly modeling. We report results on a 1,770-task evaluation slice: 1,745 Part tasks across four input regimes and RCB-Assm25, a 25-task assembly study used in every reported assembly comparison. Each method generates FreeCAD API Python, which a shared runtime executes to export the 3D model. We evaluate the exported model using executability, Solid IoU, Surface IoU, and a rubric-based visual-semantic identity Judge. Among the nine standalone frontier large models evaluated, no model leads all four metrics. Across six frontier-scale large models, executability ranges from 0.565 to 0.812, Solid IoU from 0.2841 to 0.5379, and Surface IoU from 0.112 to 0.217 across the four Part regimes. The highest regime-balanced composite comes from a different model than the leaders on the four component metrics. On RCB-Assm25, Codex with GPT-5.5 improves executability and both IoU metrics over standalone GPT-5.5, but lowers the Judge score by 6.98 percentage points, leaving GPT-5.5 as the Judge leader. We also observe recurring failure modes, most notably missing fine structures, loss of part identity, and incorrect assembly placement. These results show that execution alone is insufficient to characterize realistic CAD modeling and that frontier models and agents differ substantially across executability, IoUs, and visual-semantic identity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。