用少量数据+领域知识,让CAD模型学得更好更省数据。
KDH-CAD: Knowledge-data hybrid CAD learning under data scarcity

- 融合预训练模型与教材知识,补全设计概念
- 仅250样本达92.6%准确率,1000样本达95.8%
- 适合数据稀缺场景下的工业设计智能建模
计算机辅助设计(CAD)中的深度学习受限于数据稀缺:真实CAD数据难以大规模获取,而合成数据又难以反映实际设计实践。本文不追求更大规模的数据集,而是将CAD学习视为知识补全与校准问题。提出KDH-CAD框架,整合预训练基础模型的知识、教材/教程中的结构化领域知识以及极少量标注的CAD数据。领域知识用于挖掘并补全基础模型中表达薄弱或缺失的CAD相关概念,而少量标注数据则在隐空间校准这些概念,以适应特定任务的几何变化,无需微调基础模型。在真实机械零件分类任务上,KDH-CAD在低数据环境下表现优异:仅使用250个训练样本即达到92.6%准确率,1000样本时达95.8%,且随数据增加持续提升。性能媲美甚至超越通常需要多一个数量级数据的现有方法。结果表明,结合预训练模型与结构化领域知识可显著降低对大规模CAD数据的依赖,为高效数据驱动的CAD学习提供了可行路径。
原文摘要 · Abstract (English)
Deep learning in computer-aided design (CAD) remains fundamentally constrained by the data scarcity challenge: authentic CAD data is difficult to collect at scale, while synthetic data may not faithfully reflect real design practice. Rather than pursuing ever-larger CAD datasets, this paper alternatively treats CAD learning as a knowledge completion and calibration problem. It introduces KDH-CAD, a knowledge-data hybrid framework that integrates pretrained knowledge in foundation models, structured domain knowledge from textbooks/tutorials, and a very small amount of labeled CAD data. Domain knowledge is used to elicit and complete CAD-relevant concepts that are weakly expressed or under-represented in pretrained foundation models, while labeled CAD data calibrates these concepts in the latent space to account for task-specific geometric variability, without fine-tuning the foundation model. Experiments on real-world mechanical part classification show that KDH-CAD achieves strong performance in low-data regimes, reaching 92.6\% accuracy with only 250 training samples, 95.8\% with 1,000 samples, and continuing to improve with additional data. This matches or exceeds state-of-the-art performance that typically requires an order of magnitude more data. These results suggest that combining pretrained foundation models with structured domain knowledge can substantially reduce reliance on large-scale CAD datasets, providing a principled and practical direction for data-efficient CAD learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。