用参数+图像+文本三模态控制生成精准工程设计
Parametric-ControlNet: Multimodal Control in Foundation Models for Precise Engineering Design Synthesis
- 通过扩散模型与参数编码器实现参数化设计自动补全
- 用装配图结构整合组件图像,提升设计一致性
- 融合文本描述与多模态嵌入,支持复杂工程设计生成
本文提出一种面向工程设计合成的生成模型,可对如Stable Diffusion等文本到图像基础生成模型进行多模态控制。模型引入参数、图像和文本三种控制方式:首先,利用扩散模型与参数编码器处理部分或完整的参数输入,实现设计自动补全;其次,通过装配图系统化组装组件图像,并经组件编码器提取关键视觉信息;第三,采用CLIP编码融合文本描述,完整理解设计意图。多种输入通过多模态融合技术生成联合嵌入,作为受控于ControlNet思想的模块输入,实现对基础模型的强多模态控制,支持生成复杂且精确的工程设计。该方法拓展了AI设计工具能力,在多样化数据模态基础上显著提升设计生成的精度与多样性。
原文摘要 · Abstract (English)
This paper introduces a generative model designed for multimodal control over text-to-image foundation generative AI models such as Stable Diffusion, specifically tailored for engineering design synthesis. Our model proposes parametric, image, and text control modalities to enhance design precision and diversity. Firstly, it handles both partial and complete parametric inputs using a diffusion model that acts as a design autocomplete co-pilot, coupled with a parametric encoder to process the information. Secondly, the model utilizes assembly graphs to systematically assemble input component images, which are then processed through a component encoder to capture essential visual data. Thirdly, textual descriptions are integrated via CLIP encoding, ensuring a comprehensive interpretation of design intent. These diverse inputs are synthesized through a multimodal fusion technique, creating a joint embedding that acts as the input to a module inspired by ControlNet. This integration allows the model to apply robust multimodal control to foundation models, facilitating the generation of complex and precise engineering designs. This approach broadens the capabilities of AI-driven design tools and demonstrates significant advancements in precise control based on diverse data modalities for enhanced design generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。