用大模型结构化设计知识,高效搜索更优神经网络架构
Structuring Open-Ended NAS: Semi-Automated Design Knowledge Structuring with LLMs for Efficient Neural Architecture Search

- 用结构化模板+大模型生成多样化设计思路,突破传统搜索空间限制
- 在CIFAR-10、CIFAR-100、ImageNet16-120上分别提升0.84、2.17、2.35分
- 适合追求高效、开放型NAS探索的研究者和工程团队
当前神经网络架构搜索(NAS)方法常受限于预定义的狭窄搜索空间。尽管近期基于大语言模型(LLM)的NAS可实现开放式搜索,但往往因设计思路偏倚或质量低而效率低下。为此,我们提出半自动化结构化模型设计知识以引导搜索过程。首先定义架构属性的高层结构模板,再由LLM通过分析论文填充该模板,生成蕴含结构化设计知识的丰富多样搜索空间。为高效探索此庞大空间,我们提出FairNAD,采用多类型变异策略,包括公平采样、帕累托感知变异、LLM驱动的迭代变异及细粒度反馈循环。实验表明,FairNAD发现的架构相较现有最先进方法,在CIFAR-10、CIFAR-100和ImageNet16-120上分别提升0.84、2.17和2.35分。
原文摘要 · Abstract (English)
Current neural architecture search (NAS) methods are often limited by their predefined, restrictive search spaces. While recent large language model (LLM)-assisted NAS methods enable open-ended search spaces, they often suffer from inefficient exploration due to biased or low-quality design ideas. To address these issues, we propose to semi-automatically structure model design knowledge to guide the search process. Our approach first defines a high-level structural template of architectural attributes. An LLM then populates this template by analyzing papers, creating a rich and diverse search space that embodies this structured design knowledge. To efficiently explore this vast space, we introduce FairNAD, using a multi-type mutation that enables broad exploration through mutation with fair idea sampling, Pareto-aware mutation, LLM-driven iterative mutation, and a fine-grained feedback loop. We demonstrate the effectiveness of FairNAD in discovering high-performing architectures that yield 0.84, 2.17, and 2.35 points improvement on CIFAR-10, CIFAR-100, and ImageNet16-120, respectively, compared to current state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。