用视觉大模型+专业工具,让AI自动完成复杂建模任务
CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task Solvers
- 用视觉语言模型规划,通过Python接口调用FreeCAD执行操作
- 在多个基准上表现优于纯大模型和专用模型,支持多轮迭代优化
- 适合工业设计、自动化建模等需要跨模态理解的场景
我们提出CAD-Assistant,一个通用型人工智能辅助设计代理。该方法基于强大的视觉与大语言模型(VLLM)作为规划器,并采用针对CAD任务的工具增强范式。CAD-Assistant通过生成可在配备FreeCAD软件的Python解释器上迭代执行的动作,处理多模态用户查询。框架能够评估生成的CAD命令对几何体的影响,并根据设计状态的演变调整后续动作。我们引入多种专用于CAD的工具,包括草图图像参数化器、渲染模块、2D截面生成器及其他专用函数。在多个CAD基准上的评估显示,CAD-Assistant优于VLLM基线和监督训练的任务特定方法。此外,我们在现有基准之外,定性展示了工具增强型VLLM作为通用CAD求解器在多样化工作流中的潜力。
原文摘要 · Abstract (English)
We propose CAD-Assistant, a general-purpose CAD agent for AI-assisted design. Our approach is based on a powerful Vision and Large Language Model (VLLM) as a planner and a tool-augmentation paradigm using CAD-specific tools. CAD-Assistant addresses multimodal user queries by generating actions that are iteratively executed on a Python interpreter equipped with the FreeCAD software, accessed via its Python API. Our framework is able to assess the impact of generated CAD commands on geometry and adapts subsequent actions based on the evolving state of the CAD design. We consider a wide range of CAD-specific tools including a sketch image parameterizer, rendering modules, a 2D cross-section generator, and other specialized routines. CAD-Assistant is evaluated on multiple CAD benchmarks, where it outperforms VLLM baselines and supervised task-specific methods. Beyond existing benchmarks, we qualitatively demonstrate the potential of tool-augmented VLLMs as general-purpose CAD solvers across diverse workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。