arXiv:2506.06500cs.CL2025-06中稿 · IEEE International…被引 4

用合成数据提升大模型在电子设计中的问答能力

Improving LLM-Powered EDA Assistants with RAFT

  • 用合成问答对进行检索增强微调,弥补领域知识不足
  • 合成数据使大模型在电子设计任务中准确率显著提升
  • 适合需要安全访问与隐私保护的芯片设计团队

电子设计工程师在验证和工艺开发中常难以高效获取相关信息。尽管大语言模型(LLM)可作为对话助手提升效率,但开源预训练模型缺乏电子设计自动化(EDA)领域的专业知识。在检索增强生成(RAG)框架下,模型仍可能生成错误回答。检索增强微调(RAFT)可提升性能,但获取真实标注的问答数据在EDA领域困难。为此,我们提出使用合成问答数据来增强基于RAFT的LLM。实验表明,结合合成数据的RAFT显著提升了基于RAG的EDA任务表现。我们还研究了真实用户问题作为检索增强少样本(RAFS)示例对合成数据生成的影响。此外,系统实现了安全访问控制,确保敏感信息仅限授权人员访问。最后,评估了微调过程中数据泄露和意外记忆的风险,提供了实用建议。

原文摘要 · Abstract (English)

Electronic design engineers often struggle to efficiently access relevant information for tasks like design verification and technology development. While large language models (LLMs) can enhance productivity as conversational agents, pre-trained open-source LLMs lack domain-specific knowledge for Electronic Design Automation (EDA). In a Retrieval-Augmented Generation (RAG) context, LLMs rely on external context but may still produce inaccurate responses. Retrieval-Augmented Fine-Tuning (RAFT) improves LLM performance, but acquiring labeled question/answer (Q/A) data in EDA is difficult. To address this, we propose using synthetic Q/A datasets to enhance LLMs with RAFT. Our results show that RAFT with synthetic data significantly boosts LLM performance for RAG-based EDA tasks. We also investigate the impact of using real user questions as Retrieval-Augmented Few-Shot (RAFS) examples for synthetic data generation. Additionally, we implement secure access control to ensure sensitive information is only accessible to authorized personnel. Finally, we assess the risk of data leakage and unintended memorization during fine-tuning with synthetic data, providing practical insights.

大模型电子设计合成数据安全访问

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。