首个统一多模态多任务的CAD基准与通用模型,让一个AI搞定多种设计任务。
UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

- 用统一框架处理文本、图像、草图等多模态输入生成3D模型
- 在点云重建、图文生成、问答任务上均超越现有方法
- 适合工业设计、AI辅助建模研究者快速验证新思路
计算机辅助设计(CAD)是现代工程制造的核心,支持精确可编辑的3D模型创建。然而,当前CAD研究多聚焦单一任务,缺乏统一的多模态多任务学习基准。为此,我们提出UniCAD,一个涵盖点云到CAD重建、文本/图像到CAD生成以及CAD问答的综合性多模态基准。同时,我们构建了UniCAD-MLLM——一种通用多模态大语言模型,能端到端处理文本、图像、草图和点云输入,在单一框架内完成多种异构任务。在UniCAD和Fusion360基准上的大量实验表明,UniCAD-MLLM在所有任务中均达到领先性能,显著优于现有的专用及多任务基线。我们将开源数据集、代码与预训练模型,以加速未来研究。
原文摘要 · Abstract (English)
Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models. However, CAD research typically studies tasks in isolation, and multi-modal, multi-task learning for CAD is hindered by the absence of a unified benchmark. To address this gap, we introduce UniCAD, a comprehensive benchmark for multi-modal CAD learning that covers point-to-CAD reconstruction, text/image-to-CAD generation, and CAD question answering across diverse input modalities. Alongside the benchmark, we present UniCAD-MLLM, a universal multi-modal large language model that ingests text, images, sketches, and point clouds and performs these heterogeneous tasks in an end-to-end fashion within a single framework. Extensive experiments on the UniCAD and Fusion360 benchmarks demonstrate that UniCAD-MLLM achieves state-of-the-art performance across all tasks, outperforming existing task-specific and multi-task baselines. We will release the dataset, code, and pretrained models to accelerate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。