arXiv:2605.28579cs.AI2026-05被引 4

MUSE为文本生成CAD提供可制造、可装配的工业级评测标准

MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation

论文配图:MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation
图 1 · 摘自论文原文
  • 构建复杂装配体的结构化设计规范与三阶段评估流程
  • 强模型在工程指标上仍仅部分达标,暴露生成瓶颈
  • 适合关注工业级3D生成落地的研究者与工程师

大型语言模型(LLMs)虽推动了文本驱动3D生成,但文本到CAD仍难以支持工业产品设计。现有基准多聚焦单部件生成,以几何相似性评价,忽视功能、可制造性和可装配性。为此,我们提出MUSE,一个面向复杂可编辑边界表示(B-Rep)装配体的文本到CAD基准。MUSE将实际设计实例与结构化设计规范配对,采用三阶段评估协议:代码检查、几何检查和设计意图对齐。最终阶段使用领域专用评分标准评估功能、可制造性与可装配性,超越形状匹配,关注实际设计质量。为实现可扩展评估,采用基于评分的视觉语言模型(VLM)评判器,并通过人工标注验证其可靠性。对闭源与开源LLM的实验揭示了从可执行代码到有效几何再到工程可用设计的清晰失败链,即使最强模型在细粒度工程标准上也仅有限成功。MUSE提供了真实工业场景下的评测框架,推动文本到CAD从几何生成迈向真正工程设计。项目网站(含排行榜、数据集与代码)见https://dong7313.github.io/muse-benchmark/。

原文摘要 · Abstract (English)

Large language models (LLMs) have recently advanced text-driven 3D generation, yet Text-to-CAD remains far from supporting industrial product design. Existing benchmarks focus primarily on generating single-part CAD models and evaluate them using geometric similarity metrics that fail to capture functionality, manufacturability, and assemblability. To address this gap, we introduce MUSE, a Text-to-CAD benchmark focused on complex, editable boundary representation (B-Rep) assemblies. MUSE pairs practical design instances with structured Design Specifications and evaluates generated models through a three-stage protocol: code check, geometric check, and design-intent alignment. The final stage uses design-specific rubrics to assess functionality, manufacturability, and assemblability, moving beyond shape matching toward practical design quality. To enable scalable evaluation, we use a rubric-based visual language model (VLM) judge and validate its reliability through human annotation. Experiments on closed-source and open-source LLMs reveal a clear failure cascade from executable code to valid geometry and finally to engineering-ready design, with even the strongest models achieving limited success on fine-grained engineering criteria. Together, MUSE provides a realistic benchmark and evaluation framework for advancing Text-to-CAD from geometric generation toward true engineering design. Our project website, including the leaderboard, dataset, and code, is available at https://dong7313.github.io/muse-benchmark/.

文本生成CAD生成工业设计评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。