用强化学习提升视觉语言模型的工程制图推理能力
CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
- 分两阶段微调:先用难度感知奖励强化学习建推理能力,再监督微调指令理解
- 在20万合成+3000真实图纸上测试,显著提升复杂场景推理准确率
- 开源首个大规模制图基准数据集,适合工业设计与AI融合研究者
计算机辅助设计(CAD)在工业制造中至关重要,正交投影推理是其核心环节。现有深度学习方法多依赖标准3D重建流程,常导致尺寸不精确且限制参数化编辑。近期研究尝试使用视觉语言模型(VLMs),尤其是监督微调(SFT),但易陷入模式记忆,泛化能力差。为此,本文提出CReFT-CAD,一种两阶段微调范式:第一阶段采用课程驱动的强化学习,结合难度感知奖励逐步建立推理能力;第二阶段进行监督后微调,优化指令遵循与语义提取。同时,发布首个大规模开源基准TriView2CAD,包含20万张合成图与3000张真实世界正交投影图,附带精确尺寸标注及六种可互操作数据模态。在该基准上评估主流VLMs,结果表明CReFT-CAD显著提升推理准确率与分布外泛化能力,为推进CAD推理研究提供重要参考。
原文摘要 · Abstract (English)
Computer-Aided Design (CAD) plays a pivotal role in industrial manufacturing. Orthographic projection reasoning underpins the entire CAD workflow, encompassing design, manufacturing, and simulation. However, prevailing deep-learning approaches employ standard 3D reconstruction pipelines as an alternative, which often introduce imprecise dimensions and limit the parametric editability required for CAD workflows. Recently, some researchers adopt vision-language models (VLMs), particularly supervised fine-tuning (SFT), to tackle CAD-related challenges. SFT shows promise but often devolves into pattern memorization, yielding poor out-of-distribution performance on complex reasoning tasks. To address these gaps, we introduce CReFT-CAD, a two-stage fine-tuning paradigm that first employs a curriculum-driven reinforcement learning stage with difficulty-aware rewards to build reasoning ability steadily, and then applies supervised post-tuning to hone instruction following and semantic extraction. Complementing this, we release TriView2CAD, the first large-scale, open-source benchmark for orthographic projection reasoning, comprising 200,000 synthetic and 3,000 real-world orthographic projections with precise dimension annotations and six interoperable data modalities. We benchmark leading VLMs on orthographic projection reasoning and demonstrate that CReFT-CAD substantially improves reasoning accuracy and out-of-distribution generalizability in real-world scenarios, offering valuable insights for advancing CAD reasoning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。