arXiv:2603.13098cs.ROcs.CV2026-03中稿 · ICRA被引 3

构建24万+工业零件的多模态数据集,支持语义驱动的3D设计生成。

SldprtNet: A Large-Scale Multimodal Dataset for CAD Generation in Language-Driven 3D Design

  • 基于13种参数化命令,实现3D模型与结构化文本的无损转换。
  • 包含24.2万零件、多视图图像与自然语言描述,支持跨模态训练。
  • 适合做语义驱动建模、生成式CAD与多模态学习的研究者使用。

我们提出SldprtNet,一个包含超过242,000个工业零件的大规模数据集,用于语义驱动的CAD建模、几何深度学习及多模态模型的训练与微调。数据集提供. step和.sldprt两种格式的3D模型,支持多样化的训练与测试。为支持参数化建模并提升可扩展性,我们开发了编码器与解码器,支持13类CAD命令,并实现3D模型与结构化文本间的无损转换。每个样本配以由七个不同视角渲染图像拼接而成的复合图像,有效降低输入令牌长度并加速推理。结合编码器输出的参数化文本与复合图像,我们采用轻量级多模态模型Qwen2.5-VL-7B生成每零件的外观与功能自然语言描述。为确保准确性,我们人工验证并对齐生成描述、渲染图像与3D模型。最终,描述、参数化建模脚本、渲染图像与3D模型文件全部对齐,构成完整的SldprtNet数据集。为评估其有效性,我们在子集上微调基线模型,对比图像+文本输入与仅文本输入的表现,结果证实多模态数据集在CAD生成中的必要性与价值。该数据集涵盖真实工业零件,具备可扩展工具链、多模态内容及丰富的几何特征多样性,是专为语义驱动建模与跨模态学习构建的综合性数据集。

原文摘要 · Abstract (English)

We introduce SldprtNet, a large-scale dataset comprising over 242,000 industrial parts, designed for semantic-driven CAD modeling, geometric deep learning, and the training and fine-tuning of multimodal models for 3D design. The dataset provides 3D models in both .step and .sldprt formats to support diverse training and testing. To enable parametric modeling and facilitate dataset scalability, we developed supporting tools, an encoder and a decoder, which support 13 types of CAD commands and enable lossless transformation between 3D models and a structured text representation. Additionally, each sample is paired with a composite image created by merging seven rendered views from different viewpoints of the 3D model, effectively reducing input token length and accelerating inference. By combining this image with the parameterized text output from the encoder, we employ the lightweight multimodal language model Qwen2.5-VL-7B to generate a natural language description of each part's appearance and functionality. To ensure accuracy, we manually verified and aligned the generated descriptions, rendered images, and 3D models. These descriptions, along with the parameterized modeling scripts, rendered images, and 3D model files, are fully aligned to construct SldprtNet. To assess its effectiveness, we fine-tuned baseline models on a dataset subset, comparing image-plus-text inputs with text-only inputs. Results confirm the necessity and value of multimodal datasets for CAD generation. It features carefully selected real-world industrial parts, supporting tools for scalable dataset expansion, diverse modalities, and ensured diversity in model complexity and geometric features, making it a comprehensive multimodal dataset built for semantic-driven CAD modeling and cross-modal learning.

CAD生成多模态数据集3D设计参数化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。