arXiv:2606.05058cs.CVcs.AI2026-06被引 1

首个统一多模态多任务的CAD基准与通用模型,让一个AI搞定多种设计任务。

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

论文配图:UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD
图 1 · 摘自论文原文
  • 用统一框架处理文本、图像、草图等多模态输入生成3D模型
  • 在点云重建、图文生成、问答任务上均超越现有方法
  • 适合工业设计、AI辅助建模研究者快速验证新思路

计算机辅助设计(CAD)是现代工程制造的核心,支持精确可编辑的3D模型创建。然而,当前CAD研究多聚焦单一任务,缺乏统一的多模态多任务学习基准。为此,我们提出UniCAD,一个涵盖点云到CAD重建、文本/图像到CAD生成以及CAD问答的综合性多模态基准。同时,我们构建了UniCAD-MLLM——一种通用多模态大语言模型,能端到端处理文本、图像、草图和点云输入,在单一框架内完成多种异构任务。在UniCAD和Fusion360基准上的大量实验表明,UniCAD-MLLM在所有任务中均达到领先性能,显著优于现有的专用及多任务基线。我们将开源数据集、代码与预训练模型,以加速未来研究。

原文摘要 · Abstract (English)

Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models. However, CAD research typically studies tasks in isolation, and multi-modal, multi-task learning for CAD is hindered by the absence of a unified benchmark. To address this gap, we introduce UniCAD, a comprehensive benchmark for multi-modal CAD learning that covers point-to-CAD reconstruction, text/image-to-CAD generation, and CAD question answering across diverse input modalities. Alongside the benchmark, we present UniCAD-MLLM, a universal multi-modal large language model that ingests text, images, sketches, and point clouds and performs these heterogeneous tasks in an end-to-end fashion within a single framework. Extensive experiments on the UniCAD and Fusion360 benchmarks demonstrate that UniCAD-MLLM achieves state-of-the-art performance across all tasks, outperforming existing task-specific and multi-task baselines. We will release the dataset, code, and pretrained models to accelerate future research.

CAD生成多模态大模型工业设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。