用AI协作设计模拟电路,精准又可解释。
VLM-CAD: VLM-Optimized Collaborative Agent Design Workflow for Analog Circuit Sizing
- 将电路图转为结构化数据,让AI看懂布局。
- 在66分钟内完成设计,满足复杂指标且低功耗。
- 适合需要可靠、可解释AI辅助的芯片设计者。
视觉语言模型(VLM)在多模态推理中表现优异,但在解读密集结构化的工程内容(如模拟电路原理图)时存在空间盲区和逻辑幻觉。为此,我们提出面向模拟电路尺寸设计的视觉语言模型优化协作代理工作流(VLM-CAD),实现对多模态证据的稳健分步推理。VLM-CAD通过神经符号结构解析模块Image2Net,将原始像素转换为显式的拓扑图与结构化JSON表示,以确定性事实锚定VLM的解读。为保障工程决策可靠性,进一步提出可解释的信任区域贝叶斯优化方法ExTuRBO,利用代理生成的语义种子进行局部搜索初始化,并采用自动相关性确定法提供量化证据支持VLM决策。在两个复杂电路基准上的实验表明,VLM-CAD显著提升空间推理准确率,保持基于物理的可解释性,在满足复杂规格要求的同时实现低功耗,总运行时间低于66分钟,标志着专用技术领域中鲁棒、可解释多模态推理的重要进展。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) have demonstrated remarkable potential in multimodal reasoning, yet they inherently suffer from spatial blindness and logical hallucinations when interpreting densely structured engineering content, such as analog circuit schematics. To address these challenges, we propose a Vision Language Model-Optimized Collaborative Agent Design Workflow for Analog Circuit Sizing (VLM-CAD) designed for robust, step-by-step reasoning over multimodal evidence. VLM-CAD bridges the modality gap by integrating a neuro-symbolic structural parsing module, Image2Net, which transforms raw pixels into explicit topological graphs and structured JSON representations to anchor VLM interpretation in deterministic facts. To ensure the reliability required for engineering decisions, we further propose ExTuRBO, an Explainable Trust Region Bayesian Optimization method. ExTuRBO serves as an explainable grounding engine, employing agent-generated semantic seeds to warm-start local searches and utilizing Automatic Relevance Determination to provide quantified evidence for the VLM's decisions. Experimental results on two complex circuit benchmarks demonstrate that VLM-CAD significantly enhances spatial reasoning accuracy and maintains physics-based explainability. VLM-CAD consistently satisfies complex specification requirements while achieving low power consumption, with a total runtime under 66 minutes, marking a significant step toward robust, explainable multimodal reasoning in specialized technical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。