首个百万级工业级3D建模基准,评测AI生成可执行机械设计的能力。
OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

- 构建百万级带尺寸公差的工程图与CAD模型对,覆盖多格式数据。
- 现有模型在毫米级精度和可执行程序生成上表现仍差。
- 适合研究工业AI设计、多模态生成与智能建模的学者使用。
近期视觉语言模型(VLM)可从图像生成可执行的CAD程序,但现有方法主要针对粗粒度通用3D物体,极少关注工业机械设计所需的精细几何与毫米级公差。我们提出OmniMech,首个百万规模的多模态机械基准,用于评估VLM在工业制造数据上的可执行CAD生成能力。OmniMech包含超过25.1万张完整标注尺寸与公差的2D正交工程图,配对原始CAD模型、多视角渲染图、网格、STEP及B-rep表示,以及丰富的语义标注。基准涵盖四项任务:(1) 从工程图生成参数化CAD程序;(2) 图纸到3D的几何结构一致性重建;(3) 基于标注的尺寸、符号、特征注释与制造约束推理;(4) 利用可视化、测量、CAD执行与验证工具的工具增强型智能体推理。实验表明,当前VLM与专用CAD模型在可执行程序合成、细粒度3D重建及尺寸公差可靠执行方面仍存在显著挑战。我们将开放基准数据、评估代码与工具接口,以支持未来研究。
原文摘要 · Abstract (English)
Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D objects and rarely address the fine-grained geometry and millimeter-level tolerances required in industrial mechanical design. We introduce OmniMech, the first million-scale benchmark for evaluating VLMs on executable CAD generation from industrial manufacturing data. OmniMech contains more than 251,000 fully dimensioned and toleranced 2D orthographic drawings, paired with native CAD models, multi-view renderings, mesh, STEP and B-rep representations, and rich semantic annotations. The benchmark includes four tasks: (1) parametric CAD program synthesis from engineering drawings; (2) diagram-to-3D reasoning for geometrically and structurally consistent reconstruction; (3) annotation-grounded reasoning over dimensions, symbols, feature callouts, and manufacturing constraints; and (4) tool-augmented agentic reasoning using visualization, measurement, CAD execution, and verification tools. Experiments show that current VLMs and CAD-specialized models still struggle with executable program synthesis, fine-grained 3D reconstruction, and reliable enforcement of dimensions and tolerances. We will release the benchmark data, evaluation code, and tool interfaces to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。