arXiv:2605.10873cs.CVcs.AI2026-05被引 1

构建首个多模态CAD程序生成统一评测基准,助力AI辅助设计进步

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

论文配图:CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation
图 1 · 摘自论文原文
  • 整合六类数据集与五种输入模态,建立统一评测体系
  • 1.4万样本测试显示专业模型在理想输入下表现远超通用模型
  • 揭示复杂度、模态变化等三大失效模式,适合研究者诊断模型缺陷

从图像或3D观测中恢复可编辑的CAD程序是AI辅助设计的核心挑战,但现有评估分散在不同数据集、模态和指标间,难以衡量进展。我们提出CADBench,一个统一的多模态CAD程序生成基准。该基准包含18,000个评估样本,涵盖来自DeepCAD、Fusion 360、ABC、MCB和Objaverse的六类基准家族;支持清洁网格、带噪网格、单视角渲染、照片级渲染及多视角渲染共五种输入模态;并采用六项指标,覆盖几何保真度、可执行性与程序紧凑性。基于STEP的家族按B-rep面数分层,所有家族均通过多样性采样,支持对复杂度与对象变化的可控分析。我们对11个专用与通用视觉-语言系统进行了基准测试,生成超过140万条CAD程序。在理想输入下,专用网格到CAD模型显著优于代码生成型视觉-语言模型(VLMs),后者仍无法可靠重建CAD程序。CADBench进一步揭示三个常见失败模式:重建质量随几何复杂度下降;专用模型在模态转移下可能变得脆弱;模型排名随指标变化。这些结果使CADBench成为可编辑3D重建与多模态CAD理解进展的诊断工具。基准已开源:https://github.com/anniedoris/CADBench。

原文摘要 · Abstract (English)

Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure because existing evaluations are fragmented across datasets, modalities, and metrics. We introduce CADBench, a unified benchmark for multimodal CAD program generation. CADBench contains 18,000 evaluation samples spanning six benchmark families derived from DeepCAD, Fusion 360, ABC, MCB, and Objaverse; five input modalities including clean meshes, noisy meshes, single-view renders, photorealistic renders, and multi-view renders; and six metrics covering geometric fidelity, executability, and program compactness. STEP-based families are stratified by B-rep face count and all families are diversity-sampled to support controlled analysis across complexity and object variation. We benchmark eleven CAD-specialized and general-purpose vision-language systems, generating more than 1.4 million CAD programs. Under idealized inputs, specialized mesh-to-CAD models substantially outperform code-generating VLMs, which remain far from reliable CAD program reconstruction. CADBench further reveals three recurring failure modes: reconstruction quality degrades with geometric complexity, CAD-specialized models can be brittle under modality shift, and model rankings change across metrics. Together, these results position CADBench as a diagnostic testbed for measuring progress in editable 3D reconstruction and multimodal CAD understanding. The benchmark is publicly available at https://github.com/anniedoris/CADBench.

CAD生成多模态评测基准3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。